Skip to main content

AI May Benefit from Data Centers, but Data Centers Need Observability

Ranjan Goel
VP of Product
LogicMonitor

AI continues to shape the digital landscape and its explosion isn't slowing down anytime soon. Businesses innovate and unveil new technologies daily. In fact, we found 81% of enterprises plan to increase AI investments this year, focusing on predictive analytics, automation and anomaly detection.

This surge in AI adoption amplifies the need for robust data center infrastructure to handle the terabytes of data being generated daily. Fortunately, progress is already underway. The US government recently announced a $500 billion joint initiative in collaboration with industry leaders such as OpenAI, SoftBank, and Oracle to expand and modernize data center capabilities across the nation, ensuring the infrastructure can keep pace with AI's rapid growth.

Still, as much as AI will benefit from data centers, data centers need observability solutions to ensure resiliency and sustainability so businesses can operate to their full potential and provide seamless experiences to customers.

Why Observability Matters

Businesses have insurmountable amounts of data across IT infrastructures, and although digital transformation started over 20 years ago, many organizations are still in the process of transferring that data from on-premise solutions to the cloud, which — without the right tools in place — is a recipe for disaster of its own.

By implementing an observability solution, IT teams are given a single pane of glass view into their systems to ensure they remain up and running to reduce downtime — like the real-life scenario we saw play out with the Crowdstrike incident. With observability tools, anomalies within IT infrastructure can be detected faster, so time, resources, and money aren't lost. Coupled with next-generation AIOps tools that deliver actionable insights in order to remediate problems, observability solutions are a one-stop-shop for resilience. Multiple teams from L1 to L2 operations staff can now quickly collaborate during an incident with the same context and data all nicely summarized and root-cause identified through Agentic AI.

As IT practitioners, we know that it takes one small glitch in the system to completely flip business operations on a head, which is why these solutions are so important. Without observability, we might as well be flying blind in day-to-day operations, spending countless hours trying to rectify minor problems that cause gigantic risks. But with observability, the mean-time to resolution (MTTR) is significantly lowered allowing us to focus on mission critical work that's meaningful to our organizations at large.

Observability's Transformative Impact on Data Centers

With 68% of organizations leveraging AI tools for anomaly detection, root cause analysis, and real-time threat detection, a lot of data is being processed, and that data needs a home. Enter: data centers.

Observability comes into play to ensure those data centers remain up and running in the event of an error or software failure. If a data center were to experience an IT disruption, any system or AI that is connected to it may also fail in the process. The downtime could result in lost access to electronic records, decreased employee productivity, revenue loss, damaged customer trust and reputation, and potential compliance violations due to the service disruption.

However, the good news is that 59% of organizations that have implemented observability solutions report exceeding ROI expectations, with faster response times, improved uptime, and enhanced decision-making driving measurable business value.

Observability is a data center's best friend and it's imperative that as data centers increase in size and complexity, the investment stretches into sustainable and resilient observability solutions as well.

What's Next

The role of AI within IT operations is evolving rapidly with the advances in technology and acceptance of AI tools by operations staff. In the next 6 months, Agentic AI-driven observability and AIOps tools will become a must-have for any data center, thus improving their availability and bringing efficiency to the operations.

Ranjan Goel is VP of Product at LogicMonitor

The Latest

Two years ago, almost every customer conversation about AI started with the same questions: Which model should we use? What can it do? Is it ready for the enterprise? Today, those discussions have moved on. CIOs are far more interested in how to govern AI, integrate it with existing systems, prepare their workforce and make it part of everyday operations. The challenge is no longer to prove that AI can deliver value. It's instead about how to embed AI into the business in a way that's secure, scalable and delivers measurable outcomes ...

 

Two things happened to production incidents between 2023 and now, and they did not happen at the same speed. The first is that a class of dependency that barely existed three years ago now accounts for one incident in ten. Incidents disclosed by AI model and AI application providers rose from 1.7% of all disclosed unplanned incidents in 2023 to 10.7% in 2026 year to date, roughly a sixfold rise; that counts only incidents at AI companies themselves, so the true share is higher. The second is that the time to close an incident has not come down ...

When an AI assistant gives an incomplete or incorrect answer, teams often blame the model. They adjust prompts, switch models, increase context windows or test a new retrieval strategy. However the model may not be a problem. In many enterprise AI workflows, the problem begins inside the document-ingestion pipeline ...

If you talk to any security or observability teams right now, they're all fighting the same fire: their tooling was built to ingest X, but their sources are pumping Y and soon to be doing Z. The knee-jerk reaction is always the same: we need more platform. However, this reaction is wrong. Let me explain why, because the solution to this problem is foundational, not financial. Instead of hurling yet more money at the problem, make sure you've done what's needed upstream ...

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

AI May Benefit from Data Centers, but Data Centers Need Observability

Ranjan Goel
VP of Product
LogicMonitor

AI continues to shape the digital landscape and its explosion isn't slowing down anytime soon. Businesses innovate and unveil new technologies daily. In fact, we found 81% of enterprises plan to increase AI investments this year, focusing on predictive analytics, automation and anomaly detection.

This surge in AI adoption amplifies the need for robust data center infrastructure to handle the terabytes of data being generated daily. Fortunately, progress is already underway. The US government recently announced a $500 billion joint initiative in collaboration with industry leaders such as OpenAI, SoftBank, and Oracle to expand and modernize data center capabilities across the nation, ensuring the infrastructure can keep pace with AI's rapid growth.

Still, as much as AI will benefit from data centers, data centers need observability solutions to ensure resiliency and sustainability so businesses can operate to their full potential and provide seamless experiences to customers.

Why Observability Matters

Businesses have insurmountable amounts of data across IT infrastructures, and although digital transformation started over 20 years ago, many organizations are still in the process of transferring that data from on-premise solutions to the cloud, which — without the right tools in place — is a recipe for disaster of its own.

By implementing an observability solution, IT teams are given a single pane of glass view into their systems to ensure they remain up and running to reduce downtime — like the real-life scenario we saw play out with the Crowdstrike incident. With observability tools, anomalies within IT infrastructure can be detected faster, so time, resources, and money aren't lost. Coupled with next-generation AIOps tools that deliver actionable insights in order to remediate problems, observability solutions are a one-stop-shop for resilience. Multiple teams from L1 to L2 operations staff can now quickly collaborate during an incident with the same context and data all nicely summarized and root-cause identified through Agentic AI.

As IT practitioners, we know that it takes one small glitch in the system to completely flip business operations on a head, which is why these solutions are so important. Without observability, we might as well be flying blind in day-to-day operations, spending countless hours trying to rectify minor problems that cause gigantic risks. But with observability, the mean-time to resolution (MTTR) is significantly lowered allowing us to focus on mission critical work that's meaningful to our organizations at large.

Observability's Transformative Impact on Data Centers

With 68% of organizations leveraging AI tools for anomaly detection, root cause analysis, and real-time threat detection, a lot of data is being processed, and that data needs a home. Enter: data centers.

Observability comes into play to ensure those data centers remain up and running in the event of an error or software failure. If a data center were to experience an IT disruption, any system or AI that is connected to it may also fail in the process. The downtime could result in lost access to electronic records, decreased employee productivity, revenue loss, damaged customer trust and reputation, and potential compliance violations due to the service disruption.

However, the good news is that 59% of organizations that have implemented observability solutions report exceeding ROI expectations, with faster response times, improved uptime, and enhanced decision-making driving measurable business value.

Observability is a data center's best friend and it's imperative that as data centers increase in size and complexity, the investment stretches into sustainable and resilient observability solutions as well.

What's Next

The role of AI within IT operations is evolving rapidly with the advances in technology and acceptance of AI tools by operations staff. In the next 6 months, Agentic AI-driven observability and AIOps tools will become a must-have for any data center, thus improving their availability and bringing efficiency to the operations.

Ranjan Goel is VP of Product at LogicMonitor

The Latest

Two years ago, almost every customer conversation about AI started with the same questions: Which model should we use? What can it do? Is it ready for the enterprise? Today, those discussions have moved on. CIOs are far more interested in how to govern AI, integrate it with existing systems, prepare their workforce and make it part of everyday operations. The challenge is no longer to prove that AI can deliver value. It's instead about how to embed AI into the business in a way that's secure, scalable and delivers measurable outcomes ...

 

Two things happened to production incidents between 2023 and now, and they did not happen at the same speed. The first is that a class of dependency that barely existed three years ago now accounts for one incident in ten. Incidents disclosed by AI model and AI application providers rose from 1.7% of all disclosed unplanned incidents in 2023 to 10.7% in 2026 year to date, roughly a sixfold rise; that counts only incidents at AI companies themselves, so the true share is higher. The second is that the time to close an incident has not come down ...

When an AI assistant gives an incomplete or incorrect answer, teams often blame the model. They adjust prompts, switch models, increase context windows or test a new retrieval strategy. However the model may not be a problem. In many enterprise AI workflows, the problem begins inside the document-ingestion pipeline ...

If you talk to any security or observability teams right now, they're all fighting the same fire: their tooling was built to ingest X, but their sources are pumping Y and soon to be doing Z. The knee-jerk reaction is always the same: we need more platform. However, this reaction is wrong. Let me explain why, because the solution to this problem is foundational, not financial. Instead of hurling yet more money at the problem, make sure you've done what's needed upstream ...

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ...