Skip to main content

Seeing Is Believing - Breaking the Infrastructure Blind Spot

Tim Flower

Cloud is quickly becoming the new normal. The majority of organizations are now running at least one cloud application, and if not, they plan to do so in the near future.

The challenge for organizations is that increased cloud usage means increased complexity, often leading to a kind of infrastructure "blind spot" which puts analytic gains at risk and can obscure more pressing issues. According to a recent Forrester report, the ability to better leverage big data and analytics in business decision-making tops the priority list for organizations adopting the cloud.

So how do companies break the blind spot and get back on track?

Out of Sight?

Many companies are now adopting multiple clouds to leverage the cost-effectiveness of public resources and the granular control of private offerings. But as the Forrester research points out, choosing this route comes with multiple challenges, especially as related to infrastructure and cost visibility: 38 percent of respondents cited difficulty tracking usage across multiple clouds, while 36 percent ran into trouble monitoring costs, and 33 percent spoke to the pain point of managing network performance/latency between clouds and to/from cloud platforms.

Simply put: As cloud networks expand, so does their complexity and existing server monitoring tools aren't up to the task — they were designed to handle finite internal environments, not the ever-changing perimeter of the cloud. Under these conditions, meaningful analytics become virtually impossible since relevant data lies beyond the reach of IT observation.

Flipping the Script

The hybrid cloud breeds complexity, which limits visibility. What's the solution for multi-cloud companies that need the best of both worlds? It all starts with end-users. Think of it like this: We monitor the individual components within our data center infrastructure to make sure we maintain availability and reliability for our end users, but this legacy approach misses a whole host of external sources that can negatively impact the end user experience.

What's more, technology assessment that is strictly data-center focused automatically puts IT teams behind the curve, since end-users experiencing network problems or engaging in risky behavior — such as the use of unsanctioned cloud applications — often don't wait around for logs and error reports to reach technology pros before trying to find their own solution or downloading another app. And when they take this course of action, IT teams have minimal if non-existent methods to effectively identify this behavior and its scope across the enterprise.

As a solution, companies are turning to real user monitoring (RUM) solutions which collect data and metrics at the end-user level directly and in real-time, allowing them to effectively "flip the script" of traditional monitoring techniques. According to the Forrester survey, 77 percent of IT managers believe implementing RUM solutions would be "very effective" or "generally effective" at solving end-user monitoring challenges.

The Analytics Advantage

By adopting hybrid and multiple cloud models, businesses have access to virtually limitless data sources, but this same abundance also creates a natural "blind spot" for IT infrastructure, forcing companies to choose between reduced complexity and better analytics, or large-scale cloud adoption and limited big data effectiveness.

To make sense of all of this data, they are adopting hybrid analytics. And the emergence of flexible, RUM-based tools may suggest a way for companies to increase their visibility without losing their edge. RUM tools enable services, costs and end-users to be monitored in real-time — even as the data they provide is used to improve analytics outcomes.

As the growth of cloud continues, new advances are enabling companies to better leverage the insights gained from these multiple sources of data. Breaking the infrastructure blindspot helps remove some of the challenges of managing a new hybrid cloud-based environment.

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...

Seeing Is Believing - Breaking the Infrastructure Blind Spot

Tim Flower

Cloud is quickly becoming the new normal. The majority of organizations are now running at least one cloud application, and if not, they plan to do so in the near future.

The challenge for organizations is that increased cloud usage means increased complexity, often leading to a kind of infrastructure "blind spot" which puts analytic gains at risk and can obscure more pressing issues. According to a recent Forrester report, the ability to better leverage big data and analytics in business decision-making tops the priority list for organizations adopting the cloud.

So how do companies break the blind spot and get back on track?

Out of Sight?

Many companies are now adopting multiple clouds to leverage the cost-effectiveness of public resources and the granular control of private offerings. But as the Forrester research points out, choosing this route comes with multiple challenges, especially as related to infrastructure and cost visibility: 38 percent of respondents cited difficulty tracking usage across multiple clouds, while 36 percent ran into trouble monitoring costs, and 33 percent spoke to the pain point of managing network performance/latency between clouds and to/from cloud platforms.

Simply put: As cloud networks expand, so does their complexity and existing server monitoring tools aren't up to the task — they were designed to handle finite internal environments, not the ever-changing perimeter of the cloud. Under these conditions, meaningful analytics become virtually impossible since relevant data lies beyond the reach of IT observation.

Flipping the Script

The hybrid cloud breeds complexity, which limits visibility. What's the solution for multi-cloud companies that need the best of both worlds? It all starts with end-users. Think of it like this: We monitor the individual components within our data center infrastructure to make sure we maintain availability and reliability for our end users, but this legacy approach misses a whole host of external sources that can negatively impact the end user experience.

What's more, technology assessment that is strictly data-center focused automatically puts IT teams behind the curve, since end-users experiencing network problems or engaging in risky behavior — such as the use of unsanctioned cloud applications — often don't wait around for logs and error reports to reach technology pros before trying to find their own solution or downloading another app. And when they take this course of action, IT teams have minimal if non-existent methods to effectively identify this behavior and its scope across the enterprise.

As a solution, companies are turning to real user monitoring (RUM) solutions which collect data and metrics at the end-user level directly and in real-time, allowing them to effectively "flip the script" of traditional monitoring techniques. According to the Forrester survey, 77 percent of IT managers believe implementing RUM solutions would be "very effective" or "generally effective" at solving end-user monitoring challenges.

The Analytics Advantage

By adopting hybrid and multiple cloud models, businesses have access to virtually limitless data sources, but this same abundance also creates a natural "blind spot" for IT infrastructure, forcing companies to choose between reduced complexity and better analytics, or large-scale cloud adoption and limited big data effectiveness.

To make sense of all of this data, they are adopting hybrid analytics. And the emergence of flexible, RUM-based tools may suggest a way for companies to increase their visibility without losing their edge. RUM tools enable services, costs and end-users to be monitored in real-time — even as the data they provide is used to improve analytics outcomes.

As the growth of cloud continues, new advances are enabling companies to better leverage the insights gained from these multiple sources of data. Breaking the infrastructure blindspot helps remove some of the challenges of managing a new hybrid cloud-based environment.

The Latest

For fifteen years, observability lived downstream of everything else. Code shipped, something broke, an engineer went to the dashboards. The job was forensic. The pillars we built, such as logs, metrics, and traces, were designed for that role: tell a human what just happened, fast enough that they can make it stop. That role has quietly ended ...

Hybrid IT has become the standard operating model for enterprises — but that companies are still looking for the right hybrid IT mix, according to the 2026 State of the Data Center Report from CoreSite. After years of cloud migration and hybrid adoption, organizations are shifting their focus from deciding whether to use cloud, colocation or on-premises infrastructure to determining which workloads belong in each environment ...

Pilots are everywhere, stakeholders are seeking results, businesses are pushing for new tools, and IT teams are being asked to make AI secure, reliable, and useful at scale. But as organizations move from testing AI to operationalizing it, many are discovering that the biggest barrier is not the model, the use case, or even the budget. It is the file data foundation within ...

Fast or cheap? For most of my career in engineering, speed and quality sat on opposite ends of a seesaw. The "OR" in "fast or cheap" was non-negotiable. It was expected that pushing for faster releases meant that something in quality would give way. Tightening quality controls meant the schedule slipped. Every engineering leader I know has lived some version of that tradeoff ... The seesaw is starting to level out ...

I have been building enterprise software for more than 20 years ... One thing stays true across all of it: You do not find out your foundation is wrong during the crisis. You find out when the debt comes due. For a lot of organizations, that bill is arriving now. New research ... puts hard numbers on something practitioners have been sensing for a while. The telemetry problem isn't coming. It's already here ...

The rapid growth of AI workloads is pushing traditional log management approaches to their limits, according to The State of Log Management 2026 report from Dynatrace. Modern logs have become critical to understanding, validating, and securing AI-driven decisions, helping organizations ensure reliability, compliance, and performance at scale. However, the volume and complexity of AI telemetry are overwhelming legacy tools ...

For years, secure connectivity has relied on a familiar pattern: route traffic back to centralized gateways, inspect it, and then allow access. This model worked when applications lived in a handful of data centers and users were largely confined to offices. That model is now under strain. Applications are distributed across clouds, users connect from everywhere, and real-time workloads demand performance that centralized inspection points struggle to deliver. As traffic volumes grow and latency expectations shrink, routing everything through a small number of control points has become both a performance bottleneck and a resilience risk. The future of secure connectivity requires a different approach ...

The AI experimentation phase is over, and the private cloud is where enterprise AI workloads are being deployed for security and scale, according to Private Cloud Outlook 2026, a new report from Broadcom ... 2026 marks an acceleration into a full AI tipping point. The shift is being shaped by three forces — costs, complexity, and control — that public cloud environments are increasingly failing to address for production AI at scale. Key findings from the report include ...

44% of organizations have reported an outage in the past year tied to suppressed or ignored alerts, and 78% had at least one incident where no alert was fired at all ... Engineers learned about failures from customers. That gap between what our tools report and what our customers experience is the problem DevOps teams have been quietly solving with GenAI tooling, even as most enterprises continue to run their NOCs on manual alert triage ...

Cloud outages are usually described as technical failures. When a service goes down, a dependency breaks, or a region has issues, the focus immediately shifts to infrastructure. But if you look closely at how these incidents actually unfold, the root cause is rarely the technology itself. It is almost always tied to decisions made earlier, during design, implementation, or day-to-day operations. The system behaves the way it was built. The real question is how it was built ...