Skip to main content

How Clear Network Visibility Prevents Modern Outages

Jeremy Rossbach
Broadcom

Network engineering teams are facing a unique operational paradox: the tools designed to grant total visibility are instead creating critical data blind spots. The rapid expansion of modern enterprise tech has outpaced our ability to make sense of it.

Widespread adoption of multi-cloud environments, distributed SaaS, and microservices has caused the average digital footprint to explode over the last decade, dragging an unmanageable maze of background noise along. Today, engineers are trapped in a non-stop cycle of retrospective maintenance, buried under relentless waves of low-priority notifications, log dumps, and raw telemetry streams.

We have reached the tipping point where collecting more information no longer leads to more answers.

When a critical digital service slows, teams cannot  quickly see what is actually broken. They can't isolate the true root cause, and they have no organized way of knowing where the next outage will strike before it hits. Teams are still stuck facing the exact same fundamental unknowns. They are left trying to determine where the breakdown is, why it happened, and who is impacted.

The old infrastructure blueprint was entirely reactive. Something broke, an alert went off, and everyone rushed to fix it. Yes, this worked when everything was in a centralized data center you owned and controlled. But in a highly distributed setup, it's a fundamentally flawed system. True operational maturity doesn't start with trying to fix crashes faster; it starts with being able to be two steps ahead, cutting off the failure path before it ever reaches an end user.

Major outages are seldom random, sudden disasters.

Long before a user gets hit with a spinning wheel, the infrastructure is already showing early symptoms of a problem. It might look like a gradual capacity squeeze building up over three weeks, or brief, intermittent packet drops across an external transit path. It could be a short burst of interface errors on a local switch, or a rapid sequence of automated routing updates.

When you look at these issues in isolation, it's easy to dismiss them as routine network noise. But when you connect the dots, they tell you exactly where a crash is coming from. Organizations capable of interpreting these early signals can escape the cycle of constant triage, ensuring optimal system performance and maintaining high user satisfaction.

For years, NetOps teams focused almost entirely on filtering out low-priority alerts to prevent overload. While protecting people from endless spam matters, throwing away those minor data points is a huge mistake. An isolated glitch on a single system might look irrelevant on its own, but that exact same anomaly becomes valuable the moment you map it against your actual network topology, cloud dependencies, and active business workflows.

True operational awareness means knowing exactly how a minor disruption on an external provider echoes through your entire environment.

Observability is what changes the whole mindset and approach of an engineering organization. Transparency enables teams to investigate system health proactively, back up their architectural choices with hard evidence, and replace guesswork with clear, definitive answers.

As advanced automation and intelligent software agents are layered onto the enterprise footprint, this foundational visibility is non-negotiable. Automated systems require flawless data paths and verified environmental relationships to run safely. If there are blind spots in how these automated processes interact or where transactions are traveling across your cloud dependencies, unmitigated risks are inevitably going to hurt your business. Running automation without total visibility is taking a shot in the dark, hoping the code isn't silently disrupting operations and introducing unintended failures throughout the entire enterprise.

The ultimate competitive advantage is finding instant visibility in the infrastructure you already own.

When your services live across a messy web of external clouds and interconnected systems, staying ahead and remaining in the winning column comes down to a single capability: knowing exactly what your data is trying to tell you before undetected cracks in the foundation threaten the entire architecture.

Jeremy Rossbach is Chief Technical Evangelist - Network Observability at Broadcom

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

How Clear Network Visibility Prevents Modern Outages

Jeremy Rossbach
Broadcom

Network engineering teams are facing a unique operational paradox: the tools designed to grant total visibility are instead creating critical data blind spots. The rapid expansion of modern enterprise tech has outpaced our ability to make sense of it.

Widespread adoption of multi-cloud environments, distributed SaaS, and microservices has caused the average digital footprint to explode over the last decade, dragging an unmanageable maze of background noise along. Today, engineers are trapped in a non-stop cycle of retrospective maintenance, buried under relentless waves of low-priority notifications, log dumps, and raw telemetry streams.

We have reached the tipping point where collecting more information no longer leads to more answers.

When a critical digital service slows, teams cannot  quickly see what is actually broken. They can't isolate the true root cause, and they have no organized way of knowing where the next outage will strike before it hits. Teams are still stuck facing the exact same fundamental unknowns. They are left trying to determine where the breakdown is, why it happened, and who is impacted.

The old infrastructure blueprint was entirely reactive. Something broke, an alert went off, and everyone rushed to fix it. Yes, this worked when everything was in a centralized data center you owned and controlled. But in a highly distributed setup, it's a fundamentally flawed system. True operational maturity doesn't start with trying to fix crashes faster; it starts with being able to be two steps ahead, cutting off the failure path before it ever reaches an end user.

Major outages are seldom random, sudden disasters.

Long before a user gets hit with a spinning wheel, the infrastructure is already showing early symptoms of a problem. It might look like a gradual capacity squeeze building up over three weeks, or brief, intermittent packet drops across an external transit path. It could be a short burst of interface errors on a local switch, or a rapid sequence of automated routing updates.

When you look at these issues in isolation, it's easy to dismiss them as routine network noise. But when you connect the dots, they tell you exactly where a crash is coming from. Organizations capable of interpreting these early signals can escape the cycle of constant triage, ensuring optimal system performance and maintaining high user satisfaction.

For years, NetOps teams focused almost entirely on filtering out low-priority alerts to prevent overload. While protecting people from endless spam matters, throwing away those minor data points is a huge mistake. An isolated glitch on a single system might look irrelevant on its own, but that exact same anomaly becomes valuable the moment you map it against your actual network topology, cloud dependencies, and active business workflows.

True operational awareness means knowing exactly how a minor disruption on an external provider echoes through your entire environment.

Observability is what changes the whole mindset and approach of an engineering organization. Transparency enables teams to investigate system health proactively, back up their architectural choices with hard evidence, and replace guesswork with clear, definitive answers.

As advanced automation and intelligent software agents are layered onto the enterprise footprint, this foundational visibility is non-negotiable. Automated systems require flawless data paths and verified environmental relationships to run safely. If there are blind spots in how these automated processes interact or where transactions are traveling across your cloud dependencies, unmitigated risks are inevitably going to hurt your business. Running automation without total visibility is taking a shot in the dark, hoping the code isn't silently disrupting operations and introducing unintended failures throughout the entire enterprise.

The ultimate competitive advantage is finding instant visibility in the infrastructure you already own.

When your services live across a messy web of external clouds and interconnected systems, staying ahead and remaining in the winning column comes down to a single capability: knowing exactly what your data is trying to tell you before undetected cracks in the foundation threaten the entire architecture.

Jeremy Rossbach is Chief Technical Evangelist - Network Observability at Broadcom

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...