Skip to main content

Cloud Complexity Growing Beyond Human Abilities

Pete Goldin
APMdigest

A widening gap between IT resources and the demands of managing the increasing scale and complexity of enterprise cloud ecosystems is evident, according to Top challenges for CIOs on the road to the AI-driven autonomous cloud, a new report based on a global survey of 800 CIOs conducted by Vanson Bourne and commissioned by Dynatrace.

According to the report, IT leaders around the world are concerned about their ability to support the business effectively, as traditional monitoring solutions and custom-built approaches drown their teams in data and alerts that offer more questions than answers.

Operations teams receive nearly 3,000 alerts from their monitoring and management tools each day

CIO responses in the research indicate that, on average, IT and cloud operations teams receive nearly 3,000 alerts from their monitoring and management tools each day. With such a high volume of alerts, the average IT team spends 15% of its total available time trying to identify which alerts need to be focused on and which are irrelevant. This costs organizations an average of $1.5 million in overhead expense each year. As a result, CIOs are increasingly looking to AI and automation as they seek to maintain control and close the gap between constrained IT resources and the rising scale and complexity of the enterprise cloud.

Too Many Alerts, Not Enough Relevant Information

Traditional monitoring tools were not designed to handle the volume, velocity and variety of data generated by applications running in dynamic, web-scale enterprise clouds. These tools are often siloed and lack the broader context of events taking place across the entire technology stack. As a result, they bombard IT and cloud operations teams with hundreds, if not thousands, of alerts every day. IT is drowning in data as incremental improvements to monitoring tools fail to make a difference.

■ On average, IT and cloud operations teams receive 2,973 alerts from their monitoring and management tools each day, a 19% increase in the last 12 months.

■ 70% of CIOs say their organization is struggling to cope with the number of alerts from monitoring and management tools.

■ 75% of organizations say most of the alerts from monitoring and management tools are irrelevant.

■ On average, just 26% of the alerts organizations receive each day require action.

Traditional monitoring tools only provide data on a narrow selection of components from the technology stack. This forces IT teams to manually integrate and correlate alerts to filter out duplicates and false positives before manually identifying the underlying root cause of issues. As a result, IT teams’ ability to support the business and customers are greatly reduced as they’re faced with more questions than answers.

■ On average, IT teams spend 15% of their time trying to identify which alerts they need to focus on, and which are irrelevant.

■ The time IT teams spend trying to identify which alerts need to be focused on and which are irrelevant costs organizations, on average, $1,530,000 each year.

■ The excessive volume of alerts causes 70% of IT teams to experience problems that should have been prevented.

■ 21 incidents, on average, are experienced by organizations each year that could have been prevented if alerts were seen or acted upon in time.

■ 79% of organizations say the volume of alerts, and the time required to sift through them to identify relevant results, is making it difficult to automate enterprise cloud operations.


Methodology: This report is based on a global survey of 800 CIOs in large enterprises with over 1,000 employees, conducted by Vanson Bourne and commissioned by Dynatrace. The sample included 200 respondents in the US, 100 in the UK, France, Germany and China, and 50 in Australia, Singapore, Brazil and Mexico.

Pete Goldin is Editor and Publisher of APMdigest

Hot Topics

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Cloud Complexity Growing Beyond Human Abilities

Pete Goldin
APMdigest

A widening gap between IT resources and the demands of managing the increasing scale and complexity of enterprise cloud ecosystems is evident, according to Top challenges for CIOs on the road to the AI-driven autonomous cloud, a new report based on a global survey of 800 CIOs conducted by Vanson Bourne and commissioned by Dynatrace.

According to the report, IT leaders around the world are concerned about their ability to support the business effectively, as traditional monitoring solutions and custom-built approaches drown their teams in data and alerts that offer more questions than answers.

Operations teams receive nearly 3,000 alerts from their monitoring and management tools each day

CIO responses in the research indicate that, on average, IT and cloud operations teams receive nearly 3,000 alerts from their monitoring and management tools each day. With such a high volume of alerts, the average IT team spends 15% of its total available time trying to identify which alerts need to be focused on and which are irrelevant. This costs organizations an average of $1.5 million in overhead expense each year. As a result, CIOs are increasingly looking to AI and automation as they seek to maintain control and close the gap between constrained IT resources and the rising scale and complexity of the enterprise cloud.

Too Many Alerts, Not Enough Relevant Information

Traditional monitoring tools were not designed to handle the volume, velocity and variety of data generated by applications running in dynamic, web-scale enterprise clouds. These tools are often siloed and lack the broader context of events taking place across the entire technology stack. As a result, they bombard IT and cloud operations teams with hundreds, if not thousands, of alerts every day. IT is drowning in data as incremental improvements to monitoring tools fail to make a difference.

■ On average, IT and cloud operations teams receive 2,973 alerts from their monitoring and management tools each day, a 19% increase in the last 12 months.

■ 70% of CIOs say their organization is struggling to cope with the number of alerts from monitoring and management tools.

■ 75% of organizations say most of the alerts from monitoring and management tools are irrelevant.

■ On average, just 26% of the alerts organizations receive each day require action.

Traditional monitoring tools only provide data on a narrow selection of components from the technology stack. This forces IT teams to manually integrate and correlate alerts to filter out duplicates and false positives before manually identifying the underlying root cause of issues. As a result, IT teams’ ability to support the business and customers are greatly reduced as they’re faced with more questions than answers.

■ On average, IT teams spend 15% of their time trying to identify which alerts they need to focus on, and which are irrelevant.

■ The time IT teams spend trying to identify which alerts need to be focused on and which are irrelevant costs organizations, on average, $1,530,000 each year.

■ The excessive volume of alerts causes 70% of IT teams to experience problems that should have been prevented.

■ 21 incidents, on average, are experienced by organizations each year that could have been prevented if alerts were seen or acted upon in time.

■ 79% of organizations say the volume of alerts, and the time required to sift through them to identify relevant results, is making it difficult to automate enterprise cloud operations.


Methodology: This report is based on a global survey of 800 CIOs in large enterprises with over 1,000 employees, conducted by Vanson Bourne and commissioned by Dynatrace. The sample included 200 respondents in the US, 100 in the UK, France, Germany and China, and 50 in Australia, Singapore, Brazil and Mexico.

Pete Goldin is Editor and Publisher of APMdigest

Hot Topics

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...