Skip to main content

Why Observability Is the Missing Piece in Your Business Growth

Mimi Shalash
Splunk

For years, "observability" has been a backstage function. A quiet force that keeps digital systems running. What once lived deep in the data center is now at the center of every digital strategy. Even back in 2020, Gartner foreshadowed this shift, defining observability as "the evolution of monitoring into a process that offers insight into digital business applications, innovation, and customer experience."

That prediction has become even more relevant in an AI-driven world.

Every digital customer interaction, every cloud deployment, and every AI model depends on the same foundation: the ability to see, understand, and act on data in real time.

Digital moments like a mobile purchase, a supply chain handoff, or an AI inference run through complex layers of infrastructure. When that infrastructure falters, so does the business. Recent data from Splunk confirms that 74% of the business leaders believe observability is essential to monitoring critical business processes, and 66% feel it's key to understanding user journeys.

Because while the unknown is inevitable, observability makes it manageable. Let's explore why.

The Symbiotic Relationship Between AI and Observability

Image
Splunk

AI and observability are now inseparable. Data is the common language between them, and when used together, they amplify each other's strengths. AI helps observability teams detect patterns faster and gives engineers time back to focus on what truly moves the business forward: building better products and improving customer experiences.

Still, most teams aren't there yet. Many ITOps and engineering groups struggle with too many disconnected tools and an overload of false alerts, keeping them in a constant state of reactivity. This is the structural challenge that has persisted across enterprises: fragmented telemetry, inconsistent context, and decentralized standards.

It's no surprise, then, that organizations are turning to AI to correlate signals, reduce noise, and surface what matters most. Because prediction without observability is just speculation, and no business can afford to guess.

Splunk's research shows that 76% of practitioners now use AI regularly in daily operations, and 78% say it gives them more time to focus on innovation instead of maintenance. Yet every advancement brings new complexity. Humans in the loop are now responsible for ensuring model performance, and 47% of observability professionals say monitoring AI workloads has made their jobs more challenging, with 40% citing a lack of expertise as a barrier to AI readiness.

This gap represents a strategic opportunity. Organizations that upskill observability teams to measure AI performance and manage data quality will build a foundation of clean, governed, and trusted data that spans the entire enterprise. That means going beyond traditional IT telemetry to include operational technology (OT), IoT, sensor, and other machine data that power critical business systems.

The convergence of these once disparate data domains represents one of the most transformative opportunities in modern observability. Whether it's connecting insights from the factory floor, ERP systems, or even turbine sensors, organizations can finally uncover cross-functional intelligence that drives predictive action and measurable business outcomes.

Unlocking the Business Catalyst in Your Observability Practice

Realizing that vision requires strengthening the foundations of observability. Let's discuss four:

Minimize War Rooms and Reactivity: Many organizations still default to large, cross-functional escalations that duplicate effort and prolong mean time to resolution (MTTR). In fact, 1 in 5 respondents said they "often" or "always" start a war room that includes various departments. A more effective model emphasizes coordinated isolation and parallel response. When ITOps, engineering, and security teams visualize data through a common lens, they can trace the source of an issue faster and determine ownership. For example, when performance degradation in a key app is detected, shared telemetry allows engineering to see that the latency originates in an overloaded API gateway, not the database or underlying infrastructure. It's a simple example, but one that highlights how service mapping can triage, so people can focus on resolution, not reaction.

Get Alerting Under Control: False alerts drain engineering focus and erode trust in monitoring systems. Mature organizations address this by implementing adaptive thresholding, which dynamically adjusts alert parameters based on historical trends, system baselines, and seasonality. For example, instead of triggering dozens of CPU utilization alerts during routine batch processing every night, adaptive thresholding automatically adjusts expectations based on historical behavior. Managing alert suppression without removing early indicators of degradation is as much about data discipline as it is about process. When thresholds, alerts, and suppression logic are governed transparently and evolve with the environment, organizations build the foundation of data needed for higher levels of maturity and ultimately, AI readiness.

Lay the Foundation for Good Data That Reaps AI Benefits: Nearly half of respondents (48%) cite poor data quality as a barrier to achieving AI readiness. When engineering teams align on common data models, standardized collection practices, and comprehensive data coverage that reflects the full complexity of their environments, they establish consistent, reliable inputs for AI systems. The future belongs to organizations that can aggregate and contextualize all machine data from traditional homegrown applications, commercial off-the-shelf applications to environmental signals like temperature, vibration, and motion. When every data source speaks a common language, AI systems will be the catalyst for a new era of operational intelligence grounded in the full reality of the enterprise.

Embrace Forward Looking Architectures: The next evolution of observability is about building architectures that can adapt as fast as the systems they monitor. Organizations are investing in open and extensible technologies such as OpenTelemetry, code profiling, and observability as code to future-proof their data strategy. These approaches establish portability across environments, reduce vendor dependency, and embed observability into the software delivery lifecycle itself. OpenTelemetry, for example, is quickly becoming the industry standard for collecting, normalizing, and enriching telemetry data across hybrid and multicloud ecosystems. By adopting it early, teams can ensure consistency in how data is defined and exchanged, which sets the stage for complementary frameworks like Machine Communication Protocol (MCP). Together, these standards will be the future of advanced analytics, AI workflows, and autonomous operational systems.

Image
Splunk

Realizing Tangible Business Growth

When organizations take an innovative and responsible approach to observability, they create a foundation of agility and resilience that enables them to thrive through disruption and change. While the pace of innovation accelerates, the anchors of business success remain constant: building exceptional products, elevating customer experiences, and delivering measurable ROI that strengthens the bottom line.

In a world defined by data and driven by AI, observability is no longer just about visibility. It's now about vision.

Mimi Shalash is Observability Advisor at Splunk, a Cisco company

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Enterprise IT environments have never been more observable ... Yet many organizations still grapple with outages, lengthy incident resolution cycles, and increasing complexity. Most teams do not suffer from a shortage of data. They struggle to determine what deserves attention and what action to take next ... Enterprise IT operations must move beyond monitoring and visibility. The next stage of maturity is decision operations, an approach that helps teams make faster, better-informed decisions ...

Why Observability Is the Missing Piece in Your Business Growth

Mimi Shalash
Splunk

For years, "observability" has been a backstage function. A quiet force that keeps digital systems running. What once lived deep in the data center is now at the center of every digital strategy. Even back in 2020, Gartner foreshadowed this shift, defining observability as "the evolution of monitoring into a process that offers insight into digital business applications, innovation, and customer experience."

That prediction has become even more relevant in an AI-driven world.

Every digital customer interaction, every cloud deployment, and every AI model depends on the same foundation: the ability to see, understand, and act on data in real time.

Digital moments like a mobile purchase, a supply chain handoff, or an AI inference run through complex layers of infrastructure. When that infrastructure falters, so does the business. Recent data from Splunk confirms that 74% of the business leaders believe observability is essential to monitoring critical business processes, and 66% feel it's key to understanding user journeys.

Because while the unknown is inevitable, observability makes it manageable. Let's explore why.

The Symbiotic Relationship Between AI and Observability

Image
Splunk

AI and observability are now inseparable. Data is the common language between them, and when used together, they amplify each other's strengths. AI helps observability teams detect patterns faster and gives engineers time back to focus on what truly moves the business forward: building better products and improving customer experiences.

Still, most teams aren't there yet. Many ITOps and engineering groups struggle with too many disconnected tools and an overload of false alerts, keeping them in a constant state of reactivity. This is the structural challenge that has persisted across enterprises: fragmented telemetry, inconsistent context, and decentralized standards.

It's no surprise, then, that organizations are turning to AI to correlate signals, reduce noise, and surface what matters most. Because prediction without observability is just speculation, and no business can afford to guess.

Splunk's research shows that 76% of practitioners now use AI regularly in daily operations, and 78% say it gives them more time to focus on innovation instead of maintenance. Yet every advancement brings new complexity. Humans in the loop are now responsible for ensuring model performance, and 47% of observability professionals say monitoring AI workloads has made their jobs more challenging, with 40% citing a lack of expertise as a barrier to AI readiness.

This gap represents a strategic opportunity. Organizations that upskill observability teams to measure AI performance and manage data quality will build a foundation of clean, governed, and trusted data that spans the entire enterprise. That means going beyond traditional IT telemetry to include operational technology (OT), IoT, sensor, and other machine data that power critical business systems.

The convergence of these once disparate data domains represents one of the most transformative opportunities in modern observability. Whether it's connecting insights from the factory floor, ERP systems, or even turbine sensors, organizations can finally uncover cross-functional intelligence that drives predictive action and measurable business outcomes.

Unlocking the Business Catalyst in Your Observability Practice

Realizing that vision requires strengthening the foundations of observability. Let's discuss four:

Minimize War Rooms and Reactivity: Many organizations still default to large, cross-functional escalations that duplicate effort and prolong mean time to resolution (MTTR). In fact, 1 in 5 respondents said they "often" or "always" start a war room that includes various departments. A more effective model emphasizes coordinated isolation and parallel response. When ITOps, engineering, and security teams visualize data through a common lens, they can trace the source of an issue faster and determine ownership. For example, when performance degradation in a key app is detected, shared telemetry allows engineering to see that the latency originates in an overloaded API gateway, not the database or underlying infrastructure. It's a simple example, but one that highlights how service mapping can triage, so people can focus on resolution, not reaction.

Get Alerting Under Control: False alerts drain engineering focus and erode trust in monitoring systems. Mature organizations address this by implementing adaptive thresholding, which dynamically adjusts alert parameters based on historical trends, system baselines, and seasonality. For example, instead of triggering dozens of CPU utilization alerts during routine batch processing every night, adaptive thresholding automatically adjusts expectations based on historical behavior. Managing alert suppression without removing early indicators of degradation is as much about data discipline as it is about process. When thresholds, alerts, and suppression logic are governed transparently and evolve with the environment, organizations build the foundation of data needed for higher levels of maturity and ultimately, AI readiness.

Lay the Foundation for Good Data That Reaps AI Benefits: Nearly half of respondents (48%) cite poor data quality as a barrier to achieving AI readiness. When engineering teams align on common data models, standardized collection practices, and comprehensive data coverage that reflects the full complexity of their environments, they establish consistent, reliable inputs for AI systems. The future belongs to organizations that can aggregate and contextualize all machine data from traditional homegrown applications, commercial off-the-shelf applications to environmental signals like temperature, vibration, and motion. When every data source speaks a common language, AI systems will be the catalyst for a new era of operational intelligence grounded in the full reality of the enterprise.

Embrace Forward Looking Architectures: The next evolution of observability is about building architectures that can adapt as fast as the systems they monitor. Organizations are investing in open and extensible technologies such as OpenTelemetry, code profiling, and observability as code to future-proof their data strategy. These approaches establish portability across environments, reduce vendor dependency, and embed observability into the software delivery lifecycle itself. OpenTelemetry, for example, is quickly becoming the industry standard for collecting, normalizing, and enriching telemetry data across hybrid and multicloud ecosystems. By adopting it early, teams can ensure consistency in how data is defined and exchanged, which sets the stage for complementary frameworks like Machine Communication Protocol (MCP). Together, these standards will be the future of advanced analytics, AI workflows, and autonomous operational systems.

Image
Splunk

Realizing Tangible Business Growth

When organizations take an innovative and responsible approach to observability, they create a foundation of agility and resilience that enables them to thrive through disruption and change. While the pace of innovation accelerates, the anchors of business success remain constant: building exceptional products, elevating customer experiences, and delivering measurable ROI that strengthens the bottom line.

In a world defined by data and driven by AI, observability is no longer just about visibility. It's now about vision.

Mimi Shalash is Observability Advisor at Splunk, a Cisco company

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Enterprise IT environments have never been more observable ... Yet many organizations still grapple with outages, lengthy incident resolution cycles, and increasing complexity. Most teams do not suffer from a shortage of data. They struggle to determine what deserves attention and what action to take next ... Enterprise IT operations must move beyond monitoring and visibility. The next stage of maturity is decision operations, an approach that helps teams make faster, better-informed decisions ...