Skip to main content

5 Takeaways on the State of Observability for IT and Telecommunications Organizations

Peter Pezaris
New Relic

In May, New Relic published the State of Observability for IT and Telecommunications Report to share insights, statistics, and analysis on the adoption and business value of observability for the IT and telecommunications industries.

Here are five key takeaways from the report:

1. IT and Telco Experience More Outages than Other Industries

The report found that IT and telco organizations experience a much higher frequency of high-business-impact outages than other industries, including retail, education, healthcare, and financial services. In fact, 37% reported experiencing outages at least once a week.

Downtime costs IT and telco organizations an average of $12.71 million per year

The average annual downtime for IT and telco organizations is 26 hours, as compared to the overall average of 23 hours meaning it takes longer to detect and resolve outages. With 35% of respondents reporting that critical business app outages cost more than $500,000 per hour and 22% estimating they cost $1 million per hour, downtime costs IT and telco organizations an average of $12.71 million per year.

2. Focus on Security, Governance, Risk, and Compliance is Driving Observability Adoption

According to 50% of respondents, the top technology trend driving the need for observability amongst IT and telco organizations was an increased focus on security, governance, risk, and compliance. Despite this, the report found a diverse set of trends driving observability adoption.

Notably, compared to all industries, IT and telco respondents said some strategies and trends drive the need for observability more than others, including the development of cloud-native application architectures (48% compared to 38% overall), adoption of AI technologies (43% compared to 38% overall), and migration to a multi-cloud environment (40% compared to 37% overall).

3. Tool Consolidation is On the Rise

According to the report, IT teams in IT and telco organizations detected software and system interruptions primarily from one or more monitoring tools (77%), although about a fifth (21%) said they detect outages through manual checks or tests, complaints, or incident tickets. That said, the prevailing preference among IT and telco respondents was for a single, consolidated platform (56%). In fact, 41% said their organization is likely to consolidate tools in the next year to get the most value out of their observability spend.

The report also found that IT and telco organizations were more likely than average to use multiple monitoring tools for observability capabilities. More than two-thirds (69%) used four or more observability tools compared to 63% overall, with 23% using eight or more tools. The proportion of IT and telco respondents using a single tool has tripled since last year, growing from 1% to 3%. The data indicates that IT and telco organizations are spending time and money trying to sample and understand the different aspects of their business and avoid costly outages.

4. Observability Tooling Provides a Strong ROI

Respondents also indicated seeing high returns on observability tooling, with 57% of respondents stating their organization's return on observability investment exceeds more than $500,000 per year, including 43% who said $1 million or more. It is worth noting that IT and telco organizations reported a higher total annual value received from observability than average. These findings strongly suggest that IT and telco organizations receive a minimum 2x ROI from observability and that the ROI is even higher for organizations that monitor more of their tech stack or have a more mature observability practice.

When asked how observability helps improve their life the most, 39% of IT and Telecom decision-makers stated that observability helps establish a technology strategy, and 37% said it helps achieve technical KPIs. For practitioners, observability increases productivity so they can find and resolve issues faster (48%) and enables less guesswork when managing complicated and distributed tech stacks (32%). As far as business outcomes enabled by observability, 55% said observability improves collaboration across teams to make decisions related to the software stack. In addition, more than one-third said observability shifts developer time from incident response towards higher-value work (41%), mitigates service disruptions and business risk (37%), and improves revenue retention by deepening their understanding of customer behaviors (36%).

5. Observability Practices are Critical

It's abundantly clear that the stakes are high: If an IT or telco provider's website goes down or services are interrupted for even 30 minutes, it could cost them millions of dollars and negatively influence its customers' brand perception. Organizations simply cannot afford outages or risk losing customers due to a poor customer user experience. Data suggests that IT and telco organizations will act on their ambitious observability deployment plans in the next one to three years. By 2026, the majority of respondents expect to have deployed security monitoring (96%), network monitoring (96%), and infrastructure monitoring (94%).

Beyond preventing outages, IT and telco organizations will continue to look for ways to improve the digital customer experience, whether that is accelerating speed to market or improving uptime and reliability. While IT and telco organizations are already ahead of other industries in digital experience monitoring (DEM), even more investment in this area is anticipated as a way to create better experiences for their customers.

Organizations will continue to consolidate tools given their strong interest in deploying more capabilities in the next few years, with data indicating that more and more organizations will move from siloed tools to single, more robust platforms that provide end-to-end visibility.

Methodology: New Relic's observability forecast surveyed professionals across various industries and locations to better understand the current state of observability. Among the 1,700 technology professionals and decision-makers surveyed, 24.9% were associated with the information technology (IT) or telecommunications (telco) industries.

Peter Pezaris is Chief Design and Strategy Officer at New Relic

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

5 Takeaways on the State of Observability for IT and Telecommunications Organizations

Peter Pezaris
New Relic

In May, New Relic published the State of Observability for IT and Telecommunications Report to share insights, statistics, and analysis on the adoption and business value of observability for the IT and telecommunications industries.

Here are five key takeaways from the report:

1. IT and Telco Experience More Outages than Other Industries

The report found that IT and telco organizations experience a much higher frequency of high-business-impact outages than other industries, including retail, education, healthcare, and financial services. In fact, 37% reported experiencing outages at least once a week.

Downtime costs IT and telco organizations an average of $12.71 million per year

The average annual downtime for IT and telco organizations is 26 hours, as compared to the overall average of 23 hours meaning it takes longer to detect and resolve outages. With 35% of respondents reporting that critical business app outages cost more than $500,000 per hour and 22% estimating they cost $1 million per hour, downtime costs IT and telco organizations an average of $12.71 million per year.

2. Focus on Security, Governance, Risk, and Compliance is Driving Observability Adoption

According to 50% of respondents, the top technology trend driving the need for observability amongst IT and telco organizations was an increased focus on security, governance, risk, and compliance. Despite this, the report found a diverse set of trends driving observability adoption.

Notably, compared to all industries, IT and telco respondents said some strategies and trends drive the need for observability more than others, including the development of cloud-native application architectures (48% compared to 38% overall), adoption of AI technologies (43% compared to 38% overall), and migration to a multi-cloud environment (40% compared to 37% overall).

3. Tool Consolidation is On the Rise

According to the report, IT teams in IT and telco organizations detected software and system interruptions primarily from one or more monitoring tools (77%), although about a fifth (21%) said they detect outages through manual checks or tests, complaints, or incident tickets. That said, the prevailing preference among IT and telco respondents was for a single, consolidated platform (56%). In fact, 41% said their organization is likely to consolidate tools in the next year to get the most value out of their observability spend.

The report also found that IT and telco organizations were more likely than average to use multiple monitoring tools for observability capabilities. More than two-thirds (69%) used four or more observability tools compared to 63% overall, with 23% using eight or more tools. The proportion of IT and telco respondents using a single tool has tripled since last year, growing from 1% to 3%. The data indicates that IT and telco organizations are spending time and money trying to sample and understand the different aspects of their business and avoid costly outages.

4. Observability Tooling Provides a Strong ROI

Respondents also indicated seeing high returns on observability tooling, with 57% of respondents stating their organization's return on observability investment exceeds more than $500,000 per year, including 43% who said $1 million or more. It is worth noting that IT and telco organizations reported a higher total annual value received from observability than average. These findings strongly suggest that IT and telco organizations receive a minimum 2x ROI from observability and that the ROI is even higher for organizations that monitor more of their tech stack or have a more mature observability practice.

When asked how observability helps improve their life the most, 39% of IT and Telecom decision-makers stated that observability helps establish a technology strategy, and 37% said it helps achieve technical KPIs. For practitioners, observability increases productivity so they can find and resolve issues faster (48%) and enables less guesswork when managing complicated and distributed tech stacks (32%). As far as business outcomes enabled by observability, 55% said observability improves collaboration across teams to make decisions related to the software stack. In addition, more than one-third said observability shifts developer time from incident response towards higher-value work (41%), mitigates service disruptions and business risk (37%), and improves revenue retention by deepening their understanding of customer behaviors (36%).

5. Observability Practices are Critical

It's abundantly clear that the stakes are high: If an IT or telco provider's website goes down or services are interrupted for even 30 minutes, it could cost them millions of dollars and negatively influence its customers' brand perception. Organizations simply cannot afford outages or risk losing customers due to a poor customer user experience. Data suggests that IT and telco organizations will act on their ambitious observability deployment plans in the next one to three years. By 2026, the majority of respondents expect to have deployed security monitoring (96%), network monitoring (96%), and infrastructure monitoring (94%).

Beyond preventing outages, IT and telco organizations will continue to look for ways to improve the digital customer experience, whether that is accelerating speed to market or improving uptime and reliability. While IT and telco organizations are already ahead of other industries in digital experience monitoring (DEM), even more investment in this area is anticipated as a way to create better experiences for their customers.

Organizations will continue to consolidate tools given their strong interest in deploying more capabilities in the next few years, with data indicating that more and more organizations will move from siloed tools to single, more robust platforms that provide end-to-end visibility.

Methodology: New Relic's observability forecast surveyed professionals across various industries and locations to better understand the current state of observability. Among the 1,700 technology professionals and decision-makers surveyed, 24.9% were associated with the information technology (IT) or telecommunications (telco) industries.

Peter Pezaris is Chief Design and Strategy Officer at New Relic

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...