Skip to main content

Application Performance Management is More Than Application Performance Monitoring

Application Performance Management (APM), as defined by the industry, is focused on monitoring — because you can’t manage what you can’t see. But, there are other functions involved in managing application performance. 

For instance, this month we saw news that Outlook.com’s outage was due to a failed firmware update. Monitoring is a key element of ensuring application performance — however, other functions, such as patch management, are necessary to proactively prevent service failures. Below are a few practical considerations when delving into managing application performance.

Measuring Application Performance — What Should You Care About?

Before you start to monitor anything, you need to understand the expectations from the application’s end-users. This will help you focus on the metrics that really matter and prioritize the type of monitoring solution that is required.

For instance, is up/down monitoring adequate? Is an agentless solution sufficient? Or is something more robust needed to collect log files and so on? It’s your duty to weigh the needs of the business (i.e. what’s the impact if monitoring is not in place?) against the cost of the monitoring solution.

Having the end-user conversation will also help you understand the resource requirements for an application. Oftentimes, applications are deployed with more resources than is actually needed to meet performance objectives.

Time to Measure and Monitor — How Do You Know Application Performance is Out of Whack?

Let’s first answer this question by understanding some of the things that can go wrong:

Resources are constrained. This could happen because there is an influx of demand on the application (more users/customers). Some apps simply use more memory the longer they run. Processes can get out of control. Resource constraints can also occur if resources are shared between applications (e.g. in a virtual environment where too many VMs on the same server, SAN capacity, etc.).
 
Services stop. This can be caused by a fatal exception, etc. These things happen unexpectedly, so it’s good to have monitoring in place to alert you when a service has stopped so you can restart it immediately.

Hardware fails. Power supplies go kaput, fans break, temperature spikes, and hard drives fail. These hardware failures can and do happen, so you need advanced warning to find them and fix them quickly.

Someone changed something and it broke. Oftentimes, configuration changes can lead to performance problems. Did the Web team update the site? Was there a software update outside of a change request? Keep these peripheral factors in mind.

You’ve been hacked. According to a recent study by Ponemon Institute, survey participants experienced almost two cyber-attacks per week, many of which are DDOS attacks, as witnessed recently by Brian Krebs’ website.

Software requires updating. More often, software needs to be updated due to vulnerabilities; however, many updates fix functional bugs. In the Outlook.com example mentioned above, some functional updates can cause service outages if not applied timely and correctly.

From step 1, you have an idea of where you should focus how much of your effort. Taking it to the next step is a little tricky. For example, your application owner needs the application to be available Monday – Friday between the hours of 8 a.m. and 5 p.m., he expects no more than 1,000 users at once, and he expects users to be able to process a transaction in three minutes. 

With this information, you know critical alerts should fire during these business hours, it’s acceptable to perform software/firmware updates on the weekends or in the evening, and you have a baseline of acceptable performance from the end-user.

This application is comprised of several different components, including a Web server, application server, database and underlying hardware, storage, and networking elements. The SysAdmin is a jack of all trades who knows a little about a lot. What does it mean to monitor the SQL database? How does the SysAdmin monitor slow queries or table locks? What is a good value or a bad value? What should the threshold be? 

Luckily, there are tools that can automate a lot of the guessing and manual reporting when it comes to application performance. Tools these days should provide intelligence to what should be monitored, historical data for benchmarks/troubleshooting, and also the ability to get to the necessary details quickly.

What to Look for in Tools that Help Manage Application Performance

Application and server monitoring tools should be able to monitor across multiple components of the application to include server hardware, virtual machines, processes, services and performance metrics specific to a particular application. Tools should also provide thresholds based off best practices of what can be adjusted with historical insight as needed.

Patch management tools should provide information on which systems are out of compliance, be able to patch systems at discrete times, and inform IT when patches fail.

Configuration change management toolsshould identify and repair unauthorized configuration changes.

The time and cost associated with implementing APM tools should certainly outweigh the cost of application degradation or outage, and the IT labor costs of manually finding and fixing the problem.

ABOUT Jennifer Kuvlesky

Jennifer Kuvlesky is a Product Marketing Manager for SolarWinds, specializing in systems management. She has made her home in Austin, the high-tech capital of Texas, for more than 15 years, specializing in product management, strategy and marketing with solid knowledge of the systems and application and virtualization management market segments. Connect with Jennifer Kuvlesky on twitter @jenniferkuvlesk.

Related Links:

www.solarwinds.com

IT Budget Help: 4 Steps to Align IT Spending to Business Goals

Hot Topics

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Application Performance Management is More Than Application Performance Monitoring

Application Performance Management (APM), as defined by the industry, is focused on monitoring — because you can’t manage what you can’t see. But, there are other functions involved in managing application performance. 

For instance, this month we saw news that Outlook.com’s outage was due to a failed firmware update. Monitoring is a key element of ensuring application performance — however, other functions, such as patch management, are necessary to proactively prevent service failures. Below are a few practical considerations when delving into managing application performance.

Measuring Application Performance — What Should You Care About?

Before you start to monitor anything, you need to understand the expectations from the application’s end-users. This will help you focus on the metrics that really matter and prioritize the type of monitoring solution that is required.

For instance, is up/down monitoring adequate? Is an agentless solution sufficient? Or is something more robust needed to collect log files and so on? It’s your duty to weigh the needs of the business (i.e. what’s the impact if monitoring is not in place?) against the cost of the monitoring solution.

Having the end-user conversation will also help you understand the resource requirements for an application. Oftentimes, applications are deployed with more resources than is actually needed to meet performance objectives.

Time to Measure and Monitor — How Do You Know Application Performance is Out of Whack?

Let’s first answer this question by understanding some of the things that can go wrong:

Resources are constrained. This could happen because there is an influx of demand on the application (more users/customers). Some apps simply use more memory the longer they run. Processes can get out of control. Resource constraints can also occur if resources are shared between applications (e.g. in a virtual environment where too many VMs on the same server, SAN capacity, etc.).
 
Services stop. This can be caused by a fatal exception, etc. These things happen unexpectedly, so it’s good to have monitoring in place to alert you when a service has stopped so you can restart it immediately.

Hardware fails. Power supplies go kaput, fans break, temperature spikes, and hard drives fail. These hardware failures can and do happen, so you need advanced warning to find them and fix them quickly.

Someone changed something and it broke. Oftentimes, configuration changes can lead to performance problems. Did the Web team update the site? Was there a software update outside of a change request? Keep these peripheral factors in mind.

You’ve been hacked. According to a recent study by Ponemon Institute, survey participants experienced almost two cyber-attacks per week, many of which are DDOS attacks, as witnessed recently by Brian Krebs’ website.

Software requires updating. More often, software needs to be updated due to vulnerabilities; however, many updates fix functional bugs. In the Outlook.com example mentioned above, some functional updates can cause service outages if not applied timely and correctly.

From step 1, you have an idea of where you should focus how much of your effort. Taking it to the next step is a little tricky. For example, your application owner needs the application to be available Monday – Friday between the hours of 8 a.m. and 5 p.m., he expects no more than 1,000 users at once, and he expects users to be able to process a transaction in three minutes. 

With this information, you know critical alerts should fire during these business hours, it’s acceptable to perform software/firmware updates on the weekends or in the evening, and you have a baseline of acceptable performance from the end-user.

This application is comprised of several different components, including a Web server, application server, database and underlying hardware, storage, and networking elements. The SysAdmin is a jack of all trades who knows a little about a lot. What does it mean to monitor the SQL database? How does the SysAdmin monitor slow queries or table locks? What is a good value or a bad value? What should the threshold be? 

Luckily, there are tools that can automate a lot of the guessing and manual reporting when it comes to application performance. Tools these days should provide intelligence to what should be monitored, historical data for benchmarks/troubleshooting, and also the ability to get to the necessary details quickly.

What to Look for in Tools that Help Manage Application Performance

Application and server monitoring tools should be able to monitor across multiple components of the application to include server hardware, virtual machines, processes, services and performance metrics specific to a particular application. Tools should also provide thresholds based off best practices of what can be adjusted with historical insight as needed.

Patch management tools should provide information on which systems are out of compliance, be able to patch systems at discrete times, and inform IT when patches fail.

Configuration change management toolsshould identify and repair unauthorized configuration changes.

The time and cost associated with implementing APM tools should certainly outweigh the cost of application degradation or outage, and the IT labor costs of manually finding and fixing the problem.

ABOUT Jennifer Kuvlesky

Jennifer Kuvlesky is a Product Marketing Manager for SolarWinds, specializing in systems management. She has made her home in Austin, the high-tech capital of Texas, for more than 15 years, specializing in product management, strategy and marketing with solid knowledge of the systems and application and virtualization management market segments. Connect with Jennifer Kuvlesky on twitter @jenniferkuvlesk.

Related Links:

www.solarwinds.com

IT Budget Help: 4 Steps to Align IT Spending to Business Goals

Hot Topics

The Latest

Production incidents rarely announce themselves as database problems. They appear as slow transactions, timeouts, rising response times, or an application struggling under a workload it previously handled. APM provides an essential starting point. It can identify a slow transaction path, highlight an affected service, and show that a database dependency is consuming more time than expected. But identifying the database as part of the problem is not the same as explaining what is happening inside it ...

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments. Yet the results often disappoint ...

Ask most IT leaders about their biggest concern with AI and you'll hear the same answer: hallucinations ... Today, however, the conversation has shifted ... As organizations move beyond chatbots and experiments, they are increasingly deploying AI agents that perform multi-step tasks. These systems retrieve documents, query databases, call APIs, generate reports, write code, and make recommendations. The issue is not whether the model can reason. The issue is whether the organization can see, verify, and govern the decisions being made along the way ...

While organizations want to take control of their telemetry, building telemetry pipelines from scratch can be a very daunting, complicated task, even when leveraging open-source standards like OpenTelemetry. It requires specialized knowledge across distributed systems, data engineering, and security. This fragmented approach across systems causes higher operational costs; it puts a strain on resources and reduces efficiency as teams have to work with different interfaces and processes ...

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...