Skip to main content

Data Centers Preparing for Unpredictable AI Workload Demand

The data center industry is innovative and resilient, but also facing rising costs, worsening power constraints, and challenges in meeting the demands for AI, according to the Global Data Center Survey 2025 from Uptime Institute.

As operators expand and modernize to meet power and density requirements, they must address availability, efficiency, staffing challenges, supply chain delays, and unpredictable technological advances.

"Our data shows operators are tasked with managing a lot of big strategic challenges at the same time. These include anticipating multiple technological changes, planning for expansion in spite of major constraints on power availability, and preparing for and supporting unpredictable AI workload demand," said Andy Lawrence, Executive Director of Research, Uptime Institute. "This is a time where senior level experience is critical. But for the first time, more operators are finding it harder to recruit and retain senior people than people at an earlier stage of their career. There is a management shortage, with many experienced leaders retiring just as another phase of dramatic growth gets underway."

Roughly one-third of data center owners and operators currently perform some AI training or inference, and a significantly greater proportion plan to do so in the future. But much of this is early stage and cautious. Uncertainty over the appropriate or likely venues for AI workloads, and apprehension over the power demands of projected NVIDIA GPU systems, is likely contributing to capacity concerns.

Now in its 15th year, Uptime Institute's annual survey is the most comprehensive and longest-running study of its kind. The findings of this report highlight the practices and experiences of data center owners and operators in the areas of resiliency, sustainability, efficiency, staffing, cloud, and artificial intelligence.

Key findings from the 2025 report include:

  • Cost issues remain the top concern for digital infrastructure management teams in 2025 — but worries around forecasting future capacity requirements have grown significantly.
  • Average PUE levels show little change for the sixth consecutive year, with improvements constrained by legacy infrastructure and some climate specific limitations to efficient cooling.
  • Average server rack power densities continue to rise, with greater adoption of racks in the 10–30 kW range. Few facilities exceed 30 kW, and extreme densities are as yet rare.
  • The collection and reporting of key sustainability metrics have not improved in 2025, which is likely due in part to commercial pressures to support AI, and easing regulatory pressure in some regions.
  • Trust in AI for data center operations depends on the use case: most would allow its use for analyzing sensor data and predictive maintenance tasks, but not configuration changes, controlling equipment, or staffing issues.
  • Impactful data center outages are gradually becoming less frequent — but one in ten still cause serious or severe disruption, underscoring the need for continued investment.
  • Enterprises continue to adopt hybrid IT strategies, spanning cloud, colocation on-premises data centers. On-premises data centers remain foundational for those with large, mission critical processing needs, with 45% of IT workloads still residing in corporate facilities.
  • Staffing challenges persist in 2025. Nearly two-thirds of operators report difficulty retaining staff, finding qualified candidates, or both.

Methodology: Uptime conducted this year's Annual Global Data Center Survey online and via email from April to May 2025 and collected responses from more than 800 data center owners and operators. For the third consecutive year, Uptime's survey asked data center operators to identify their management team's top concerns related to digital infrastructure. In 2025, new response options were added to reflect the evolving challenges surrounding power availability, supply chain disruptions, and demand for AI.

The survey participants represent a wide range of industry verticals in multiple countries. Nearly half (43%) are located in North America and Europe. Approximately one in five respondents work for professional IT / data center service providers — that is, staff with operational or executive responsibilities for a third-party data center, such as those offering colocation, wholesale, software or cloud computing services.

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Enterprise IT environments have never been more observable ... Yet many organizations still grapple with outages, lengthy incident resolution cycles, and increasing complexity. Most teams do not suffer from a shortage of data. They struggle to determine what deserves attention and what action to take next ... Enterprise IT operations must move beyond monitoring and visibility. The next stage of maturity is decision operations, an approach that helps teams make faster, better-informed decisions ...

Data Centers Preparing for Unpredictable AI Workload Demand

The data center industry is innovative and resilient, but also facing rising costs, worsening power constraints, and challenges in meeting the demands for AI, according to the Global Data Center Survey 2025 from Uptime Institute.

As operators expand and modernize to meet power and density requirements, they must address availability, efficiency, staffing challenges, supply chain delays, and unpredictable technological advances.

"Our data shows operators are tasked with managing a lot of big strategic challenges at the same time. These include anticipating multiple technological changes, planning for expansion in spite of major constraints on power availability, and preparing for and supporting unpredictable AI workload demand," said Andy Lawrence, Executive Director of Research, Uptime Institute. "This is a time where senior level experience is critical. But for the first time, more operators are finding it harder to recruit and retain senior people than people at an earlier stage of their career. There is a management shortage, with many experienced leaders retiring just as another phase of dramatic growth gets underway."

Roughly one-third of data center owners and operators currently perform some AI training or inference, and a significantly greater proportion plan to do so in the future. But much of this is early stage and cautious. Uncertainty over the appropriate or likely venues for AI workloads, and apprehension over the power demands of projected NVIDIA GPU systems, is likely contributing to capacity concerns.

Now in its 15th year, Uptime Institute's annual survey is the most comprehensive and longest-running study of its kind. The findings of this report highlight the practices and experiences of data center owners and operators in the areas of resiliency, sustainability, efficiency, staffing, cloud, and artificial intelligence.

Key findings from the 2025 report include:

  • Cost issues remain the top concern for digital infrastructure management teams in 2025 — but worries around forecasting future capacity requirements have grown significantly.
  • Average PUE levels show little change for the sixth consecutive year, with improvements constrained by legacy infrastructure and some climate specific limitations to efficient cooling.
  • Average server rack power densities continue to rise, with greater adoption of racks in the 10–30 kW range. Few facilities exceed 30 kW, and extreme densities are as yet rare.
  • The collection and reporting of key sustainability metrics have not improved in 2025, which is likely due in part to commercial pressures to support AI, and easing regulatory pressure in some regions.
  • Trust in AI for data center operations depends on the use case: most would allow its use for analyzing sensor data and predictive maintenance tasks, but not configuration changes, controlling equipment, or staffing issues.
  • Impactful data center outages are gradually becoming less frequent — but one in ten still cause serious or severe disruption, underscoring the need for continued investment.
  • Enterprises continue to adopt hybrid IT strategies, spanning cloud, colocation on-premises data centers. On-premises data centers remain foundational for those with large, mission critical processing needs, with 45% of IT workloads still residing in corporate facilities.
  • Staffing challenges persist in 2025. Nearly two-thirds of operators report difficulty retaining staff, finding qualified candidates, or both.

Methodology: Uptime conducted this year's Annual Global Data Center Survey online and via email from April to May 2025 and collected responses from more than 800 data center owners and operators. For the third consecutive year, Uptime's survey asked data center operators to identify their management team's top concerns related to digital infrastructure. In 2025, new response options were added to reflect the evolving challenges surrounding power availability, supply chain disruptions, and demand for AI.

The survey participants represent a wide range of industry verticals in multiple countries. Nearly half (43%) are located in North America and Europe. Approximately one in five respondents work for professional IT / data center service providers — that is, staff with operational or executive responsibilities for a third-party data center, such as those offering colocation, wholesale, software or cloud computing services.

The Latest

For decades, enterprise networks were designed around a simple assumption: work happened inside the office. Applications lived in centralized data centers, employees connected through internal infrastructure, and security focused on protecting the perimeter that surrounded everything ... But the way organizations operate today bears little resemblance to that environment. Cloud platforms host critical applications, employees connect from homes and airports as often as they do from offices, and partners collaborate through shared systems that exist far beyond corporate walls. In short, the corporate network no longer resembles the environment it was designed to protect ...

As an analyst who researches how IT organizations design, build, and operate their networks, I find that network data is a constant source of pain. Network teams struggle with data quality, fragmentation, authority, access, and trust. And these issues undermine everything they try to do. Here are the numbers: Only 45% of network teams are completely confident in the accuracy of their network source of truth, which documents the intent of their network ...

The 2026 Global Data Center Survey from Uptime Institute reveals an industry navigating workforce constraints, escalating outage expenses, even as rising costs remain the top concern for management teams ...

The next observability gap may not be in the code. It may be under the rack. That sounds strange until you think about how AI incidents actually feel in the middle of an investigation ... The application dashboard may be accurate. It may also be stopping at the wrong boundary. AI systems depend on software, but they also depend on a dense physical stack: racks, power paths, thermal margin, maintenance activity and, in many environments, liquid cooling. Those physical dependencies can change slowly before they look like a software incident ...

Certificate expiration is the rare outage you can see coming. Every TLS certificate carries the date it stops working, so the moment it will begin breaking connections is knowable in advance. That's what makes an expired certificate such a frustrating way to lose a service. What's changing now is how often that date comes around ...

Enterprises operate different combinations of workloads across cloud, hybrid and multicloud environments. For business-critical workloads, teams need to consider monitoring and observability early so they can detect health issues, investigate failures, and understand operational impact. Organizations place workloads on cloud platforms based on a combination of technical requirements, economics, existing dependencies, organizational standards, and business priorities. Their monitoring priorities therefore depend on what they operate and where those systems run. Those priorities will not look the same for every organization ...

Top-performing businesses prioritize data-driven decision making, enabling leaders to move from intuition and gut feel towards evidence-based judgment. But that judgment is only sound when the data underpinning decisions is accurate. With incident management, data accuracy is particularly important. Long-term revenue, customer trust, and operational stability depend on high-quality data that enables teams to quickly identify and address the root cause of major incidents. Against this backdrop, governance becomes a critical endeavor to ensure the right data drives the right action ...

In MEAN TIME TO INSIGHT Episode 26, Shamus McGillicuddy, VP of Research, Network Infrastructure and Operations, at EMA discusses network compliance ... 

Most production autonomous agents do not run in a vacuum. They run inside cloud infrastructure: virtual machines, containers, pods, managed clusters or private servers. That is where most operations teams start monitoring. Is the VM alive? Is the container running? Did the pod restart? Is memory stable? Is CPU too high? Did the health check pass? Those signals are useful. They tell you whether the shell around the agent is alive. They do not tell you whether the agent inside is actually operational ...

Enterprise IT environments have never been more observable ... Yet many organizations still grapple with outages, lengthy incident resolution cycles, and increasing complexity. Most teams do not suffer from a shortage of data. They struggle to determine what deserves attention and what action to take next ... Enterprise IT operations must move beyond monitoring and visibility. The next stage of maturity is decision operations, an approach that helps teams make faster, better-informed decisions ...