Skip to main content

NetOps Teams: Hybrid Cloud Architecture Is the True Source of Our Pain, Not the Cloud Itself

Shamus McGillicuddy

Hybrid cloud architecture is breaking the backs of network engineering and operations teams. These teams are more successful when their companies go all-in with the cloud or stay out of it entirely. When companies maintain hybrid infrastructure, with applications and data residing across data centers and public cloud services, the network team struggles.

This insight emerged in the newly published 2024 edition of Enterprise Management Associates' (EMA) Network Management Megatrends research, our biennial network operations benchmarking study based on a survey of 406 enterprise IT professionals. The survey found that nearly 51% of respondents were supporting a hybrid cloud environment. More than 26% were 100% in the public cloud, while more than 23% were relying entirely on private data center infrastructure.

EMA has long observed that network operations teams are challenged by the public cloud. Network engineers tell us that they lose control over architecture and struggle to gain visibility into cloud networks. They also struggle with skills gaps and collaboration challenges. However, this most recent finding suggests that it isn't just the cloud itself that necessarily causes these problems. Instead, it's the hybridization of infrastructure, with resources spanning public and private resources, that is causing significant pain. Network teams are finding it difficult to manage across the two worlds. They lack tools and processes for architecting and operationalizing an effective hybrid cloud network. To illustrate this point, the Megatrends research asked respondents to identify the typical root cause of complex service issues that pull together war room response teams with stakeholders from multiple IT groups. In hybrid cloud enterprises, network infrastructure was more likely to be the root cause of these challenging problems.

When one thinks about it, the challenge of a hybrid environment makes perfect sense. In a company that uses only private infrastructure, network teams have long had the tools, processes, and skills required to manage such networks. Those that adopt a 100% cloud environment can adapt by embracing cloud-native technologies for building, managing and observing networks. There are countless vendors that have developed solutions that excel at supporting purely cloud networks, and cloud providers themselves have built out an ecosystem of tools to enable networks within their environments.

It's much harder to build and manage networks that span these two worlds because they are fundamentally different. In private infrastructure, the network team has administrative access and control over the underlying infrastructure, making it easier to manage networks granularly and to extract monitoring data. In the public cloud, network teams must work with APIs and abstracted services because they have no administrative access to the underlying infrastructure. Deploying tools and developing processes that span these very different architectures is fundamentally complex.

Addressing Hybrid Cloud Complexity

The Megatrends research revealed that network teams that support a hybrid cloud architecture are more likely to be investing in new network performance management tools, suggesting that they've identified gaps in how their existing tools provide end-to-end visibility into hybrid cloud networks. They are looking for tools that can provide end to end visibility. In fact, 53% of network teams who support a hybrid cloud or multi-cloud environment reported that they are trying to build out end-to-end monitoring and troubleshooting capabilities. Network teams have a variety of options here, from their legacy on-premises tools and cloud-native observability tools, to tools that focus more on digital experience monitoring than infrastructure monitoring. It remains unclear which approach is best.

"I see a lack of holistic visibility into the cloud. It's hard to have a single pane of glass when you are looking at various systems," as a network security architect at a Fortune 500 cybersecurity company recently told EMA. "We've been relying on logs. But there are not tools that are designed to give us visibility across different [public and private] clouds."

Network teams with hybrid cloud environments are also more likely to be investing in new network security solutions, suggesting that they are trying to impose consistent and effective security controls across public and private infrastructure. These network teams were also more likely to share their network management and monitoring tools with their counterparts in cybersecurity teams, suggesting that the complexity of hybrid cloud security forces more collaboration between network and security teams.

Hybrid cloud network teams are also more likely to make new investments in multi-cloud networking solutions, software-defined WAN, and data center network infrastructure, indicating that hybrid environments are demanding modernization of local and wide-area network connectivity.

In other words, hybrid cloud infrastructure simply demands more of network teams than pure private and pure cloud environments. The complexity doesn't end there. EMA also observed that companies with hybrid environments are more likely to have a mix of on-premises data centers and colocation providers comprising their private infrastructure. Research participants told us that they typically incorporate colocation data centers into their networks to support efforts to reduce their on-premises data center footprints, to enable disaster recovery and resiliency, and to deploy applications and data closer to end users and customers.

Finally, all this complexity creates a need for operational efficiency. One key enabler of that efficiency is network automation. Unfortunately, enterprises with hybrid environments were more likely than others to struggle to hire IT personnel with network automation skills. Network teams will need vendors to support them as they try to mitigate these skills gaps.

Overall, just 39% of network pros who support a hybrid cloud and/or multi-cloud architecture believe that their efforts to impose end-to-end operations have been completely successful. EMA observed two potential paths toward success:

First, network teams that have improved their ability to hire and retain skilled personnel do better with these efforts. This will require IT leadership that is willing to go the extra mile, with improved recruiting efforts, competitive compensation, and opportunities for advancement and growth.

Second, network teams do better when they minimize their reliance on tools provided by individual cloud providers to operationalize hybrid networks. Instead, they need to push their traditional tool vendors to innovate.

Listen to Episode 1 of the MEAN TIME TO INSIGHT PODCAST (MTTI): 2024 Network Management Trends

Click here for a direct MP3 download of Episode 1

Hot Topics

The Latest

Performance bottlenecks aren't uncommon when it comes to rolling out new technology, regardless of how capable or game-changing that technology might be. Every generation of new tech has encountered roadblocks that had to be overcome before it was truly able to shine. Virtualization forced organizations to rethink resource allocation, cloud transformation had us shift our focus toward scalability and elasticity, and microservices introduced entirely new challenges around observability and distributed systems. There's something different about AI, however ...

Consider a single order represented across order-management, execution, and settlement systems. Each database, message broker, and application may be online and processing its own records correctly. Yet the workflow has failed if related events arrive on different clocks, rely on inconsistent state, or cannot be reconciled before an operational decision must be made ...

AI now exists in almost every IT workflow. In a recent survey of more than 800 IT service professionals, all respondents indicated the use of AI in some form within their organization. But there's a growing paradox: if dashboards are clearing faster and alerts are resolved at unprecedented speed, why aren't IT service desks reporting lighter workloads? The research found that 71% of IT teams said their actual workload has remained flat or increased since adopting AI. This reality appears to contradict what we’ve been told about AI ...

Two years ago, almost every customer conversation about AI started with the same questions: Which model should we use? What can it do? Is it ready for the enterprise? Today, those discussions have moved on. CIOs are far more interested in how to govern AI, integrate it with existing systems, prepare their workforce and make it part of everyday operations. The challenge is no longer to prove that AI can deliver value. It's instead about how to embed AI into the business in a way that's secure, scalable and delivers measurable outcomes ...

Two things happened to production incidents between 2023 and now, and they did not happen at the same speed. The first is that a class of dependency that barely existed three years ago now accounts for one incident in ten. Incidents disclosed by AI model and AI application providers rose from 1.7% of all disclosed unplanned incidents in 2023 to 10.7% in 2026 year to date, roughly a sixfold rise; that counts only incidents at AI companies themselves, so the true share is higher. The second is that the time to close an incident has not come down ...

When an AI assistant gives an incomplete or incorrect answer, teams often blame the model. They adjust prompts, switch models, increase context windows or test a new retrieval strategy. However the model may not be a problem. In many enterprise AI workflows, the problem begins inside the document-ingestion pipeline ...

If you talk to any security or observability teams right now, they're all fighting the same fire: their tooling was built to ingest X, but their sources are pumping Y and soon to be doing Z. The knee-jerk reaction is always the same: we need more platform. However, this reaction is wrong. Let me explain why, because the solution to this problem is foundational, not financial. Instead of hurling yet more money at the problem, make sure you've done what's needed upstream ...

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

NetOps Teams: Hybrid Cloud Architecture Is the True Source of Our Pain, Not the Cloud Itself

Shamus McGillicuddy

Hybrid cloud architecture is breaking the backs of network engineering and operations teams. These teams are more successful when their companies go all-in with the cloud or stay out of it entirely. When companies maintain hybrid infrastructure, with applications and data residing across data centers and public cloud services, the network team struggles.

This insight emerged in the newly published 2024 edition of Enterprise Management Associates' (EMA) Network Management Megatrends research, our biennial network operations benchmarking study based on a survey of 406 enterprise IT professionals. The survey found that nearly 51% of respondents were supporting a hybrid cloud environment. More than 26% were 100% in the public cloud, while more than 23% were relying entirely on private data center infrastructure.

EMA has long observed that network operations teams are challenged by the public cloud. Network engineers tell us that they lose control over architecture and struggle to gain visibility into cloud networks. They also struggle with skills gaps and collaboration challenges. However, this most recent finding suggests that it isn't just the cloud itself that necessarily causes these problems. Instead, it's the hybridization of infrastructure, with resources spanning public and private resources, that is causing significant pain. Network teams are finding it difficult to manage across the two worlds. They lack tools and processes for architecting and operationalizing an effective hybrid cloud network. To illustrate this point, the Megatrends research asked respondents to identify the typical root cause of complex service issues that pull together war room response teams with stakeholders from multiple IT groups. In hybrid cloud enterprises, network infrastructure was more likely to be the root cause of these challenging problems.

When one thinks about it, the challenge of a hybrid environment makes perfect sense. In a company that uses only private infrastructure, network teams have long had the tools, processes, and skills required to manage such networks. Those that adopt a 100% cloud environment can adapt by embracing cloud-native technologies for building, managing and observing networks. There are countless vendors that have developed solutions that excel at supporting purely cloud networks, and cloud providers themselves have built out an ecosystem of tools to enable networks within their environments.

It's much harder to build and manage networks that span these two worlds because they are fundamentally different. In private infrastructure, the network team has administrative access and control over the underlying infrastructure, making it easier to manage networks granularly and to extract monitoring data. In the public cloud, network teams must work with APIs and abstracted services because they have no administrative access to the underlying infrastructure. Deploying tools and developing processes that span these very different architectures is fundamentally complex.

Addressing Hybrid Cloud Complexity

The Megatrends research revealed that network teams that support a hybrid cloud architecture are more likely to be investing in new network performance management tools, suggesting that they've identified gaps in how their existing tools provide end-to-end visibility into hybrid cloud networks. They are looking for tools that can provide end to end visibility. In fact, 53% of network teams who support a hybrid cloud or multi-cloud environment reported that they are trying to build out end-to-end monitoring and troubleshooting capabilities. Network teams have a variety of options here, from their legacy on-premises tools and cloud-native observability tools, to tools that focus more on digital experience monitoring than infrastructure monitoring. It remains unclear which approach is best.

"I see a lack of holistic visibility into the cloud. It's hard to have a single pane of glass when you are looking at various systems," as a network security architect at a Fortune 500 cybersecurity company recently told EMA. "We've been relying on logs. But there are not tools that are designed to give us visibility across different [public and private] clouds."

Network teams with hybrid cloud environments are also more likely to be investing in new network security solutions, suggesting that they are trying to impose consistent and effective security controls across public and private infrastructure. These network teams were also more likely to share their network management and monitoring tools with their counterparts in cybersecurity teams, suggesting that the complexity of hybrid cloud security forces more collaboration between network and security teams.

Hybrid cloud network teams are also more likely to make new investments in multi-cloud networking solutions, software-defined WAN, and data center network infrastructure, indicating that hybrid environments are demanding modernization of local and wide-area network connectivity.

In other words, hybrid cloud infrastructure simply demands more of network teams than pure private and pure cloud environments. The complexity doesn't end there. EMA also observed that companies with hybrid environments are more likely to have a mix of on-premises data centers and colocation providers comprising their private infrastructure. Research participants told us that they typically incorporate colocation data centers into their networks to support efforts to reduce their on-premises data center footprints, to enable disaster recovery and resiliency, and to deploy applications and data closer to end users and customers.

Finally, all this complexity creates a need for operational efficiency. One key enabler of that efficiency is network automation. Unfortunately, enterprises with hybrid environments were more likely than others to struggle to hire IT personnel with network automation skills. Network teams will need vendors to support them as they try to mitigate these skills gaps.

Overall, just 39% of network pros who support a hybrid cloud and/or multi-cloud architecture believe that their efforts to impose end-to-end operations have been completely successful. EMA observed two potential paths toward success:

First, network teams that have improved their ability to hire and retain skilled personnel do better with these efforts. This will require IT leadership that is willing to go the extra mile, with improved recruiting efforts, competitive compensation, and opportunities for advancement and growth.

Second, network teams do better when they minimize their reliance on tools provided by individual cloud providers to operationalize hybrid networks. Instead, they need to push their traditional tool vendors to innovate.

Listen to Episode 1 of the MEAN TIME TO INSIGHT PODCAST (MTTI): 2024 Network Management Trends

Click here for a direct MP3 download of Episode 1

Hot Topics

The Latest

Performance bottlenecks aren't uncommon when it comes to rolling out new technology, regardless of how capable or game-changing that technology might be. Every generation of new tech has encountered roadblocks that had to be overcome before it was truly able to shine. Virtualization forced organizations to rethink resource allocation, cloud transformation had us shift our focus toward scalability and elasticity, and microservices introduced entirely new challenges around observability and distributed systems. There's something different about AI, however ...

Consider a single order represented across order-management, execution, and settlement systems. Each database, message broker, and application may be online and processing its own records correctly. Yet the workflow has failed if related events arrive on different clocks, rely on inconsistent state, or cannot be reconciled before an operational decision must be made ...

AI now exists in almost every IT workflow. In a recent survey of more than 800 IT service professionals, all respondents indicated the use of AI in some form within their organization. But there's a growing paradox: if dashboards are clearing faster and alerts are resolved at unprecedented speed, why aren't IT service desks reporting lighter workloads? The research found that 71% of IT teams said their actual workload has remained flat or increased since adopting AI. This reality appears to contradict what we’ve been told about AI ...

Two years ago, almost every customer conversation about AI started with the same questions: Which model should we use? What can it do? Is it ready for the enterprise? Today, those discussions have moved on. CIOs are far more interested in how to govern AI, integrate it with existing systems, prepare their workforce and make it part of everyday operations. The challenge is no longer to prove that AI can deliver value. It's instead about how to embed AI into the business in a way that's secure, scalable and delivers measurable outcomes ...

Two things happened to production incidents between 2023 and now, and they did not happen at the same speed. The first is that a class of dependency that barely existed three years ago now accounts for one incident in ten. Incidents disclosed by AI model and AI application providers rose from 1.7% of all disclosed unplanned incidents in 2023 to 10.7% in 2026 year to date, roughly a sixfold rise; that counts only incidents at AI companies themselves, so the true share is higher. The second is that the time to close an incident has not come down ...

When an AI assistant gives an incomplete or incorrect answer, teams often blame the model. They adjust prompts, switch models, increase context windows or test a new retrieval strategy. However the model may not be a problem. In many enterprise AI workflows, the problem begins inside the document-ingestion pipeline ...

If you talk to any security or observability teams right now, they're all fighting the same fire: their tooling was built to ingest X, but their sources are pumping Y and soon to be doing Z. The knee-jerk reaction is always the same: we need more platform. However, this reaction is wrong. Let me explain why, because the solution to this problem is foundational, not financial. Instead of hurling yet more money at the problem, make sure you've done what's needed upstream ...

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...