Skip to main content

From Cleanup to Prevention: The Future of Cloud-Native Cost Optimization

Adi Fayer
Komodor

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments.

Yet the results often disappoint. The first wave of savings comes quickly, then progress stalls. Clusters still run larger than they should, while nodes remain partially empty but cannot be removed. Platform teams see idle capacity, but the infrastructure cannot safely release it. The finance department sees ongoing waste, but engineering teams insist the environment has already been optimized.

Both sides are right.

The next phase of savings will come from preventing inefficient cluster states before they form.

Why Reactive Optimization Hits a Ceiling

Most cost programs start after waste already exists. Rightsizing tools adjust workload requests. Autoscalers try to consolidate nodes. FinOps platforms show where spend is high. These capabilities are useful, but they operate on a cluster state that has already been shaped by earlier scheduling decisions.

That is the core limitation.

The Kubernetes scheduler places pods based on what fits at the moment. Autoscalers try to clean things up later by removing capacity that is no longer needed. But if workloads have already landed in ways that scatter unevictable pods, create fragmentation, or violate consolidation paths, the autoscaler inherits a problem it isn't designed to solve.

This explains why teams using Karpenter or Cluster Autoscaler may still see significant capacity remain idle or locked.

The Real Problem Is Cluster State

The problem with reactive optimization is that it often acts on waste that is already entrenched and has become difficult to remove.

By the time a dashboard shows unused CPU or memory, that capacity may already be scattered across nodes the cluster cannot safely drain. Some nodes may be pinned by configurations associated with Pod Disruption Budgets or affinity rules. On paper, the cluster has available capacity, but that capacity cannot be practically consolidated into fewer running machines.

This is why utilization improvements have a ceiling. Rightsizing, autoscaler tuning, and better dashboards can only act on the cluster state that already exists. They may expose unused capacity, but they cannot always undo the placement decisions, policies, and workload constraints that prevent nodes from being drained. When that happens, the waste remains structurally locked into the cluster.

Dashboards may show improvement. Teams may feel they have done the right work. But if inefficiently utilized nodes cannot be drained and terminated, infrastructure spend does not fall.

That's why the cost conversation has to move beyond how much capacity each workload requests and focus on whether the cluster can be consolidated using existing infrastructure to reduce idle resources.

This requires a different operating model.

Move Optimization Earlier

The most important shift is pushing cost intelligence closer to scheduling time.

Instead of waiting for the cluster to fragment and then trying to consolidate it later, platform teams need to influence placement before workloads land. That means evaluating not only whether a pod fits now, but whether placing it on a particular node will make the cluster harder to scale down later.

This is a proactive scaling mindset. It considers workload behavior, reliability constraints, eviction rules, autoscaler logic, and future drain scenarios before placement decisions create waste. If a node is likely to be removed, the platform should avoid deploying new workloads there. If certain pods are difficult to evict, they should be placed in dedicated nodes rather than scattered randomly. If complementary workloads can share capacity efficiently, the platform should account for that before idle resources become stranded.

The goal should be to keep the environment flexible enough for autoscalers to do their job more efficiently.

Detection and Prevention Must Work Together

Proactive placement is only part of the answer. Most enterprise clusters already contain years of accumulated constraints, exceptions, and workload patterns that block consolidation. Preventing new waste does not automatically eliminate old waste.

That means the modern cost model needs two loops.

The first loop detects existing blockers: workloads that prevent node removal, overly restrictive disruption policies, autoscaler settings that preserve waste, and nodes kept alive by a small number of unevictable pods. The second loop prevents new inefficiency by guiding placement decisions before those blockers spread.

When these loops work together, cost optimization becomes continuous state management. The platform is not merely identifying waste after the fact. It is keeping the cluster in a condition where savings can actually be realized.

This is a more mature view of cloud-native efficiency. It recognizes that the best cost decision is not always the cheapest immediate placement. It is the placement that preserves reliability while allowing infrastructure to scale down cleanly over time.

From Cleanup to Cost-Aware Design

The next layer of cloud-native savings will not come from pushing teams to rightsize harder. That work still matters, but it is not enough.

The larger opportunity is to engineer efficiency into the platform before waste takes hold. That requires treating scheduling, autoscaling, and reliability policy as part of the same cost system.

Cloud-native infrastructure was built to be dynamic, but many cost practices remain static and reactive. They focus on waste only after it appears, explain why spend is high after the bill arrives, then expect autoscalers to clean up placement decisions that were never optimized for future consolidation.

That model is reaching its limit. The next phase of cloud cost management needs to be proactive. It should prevent bad cluster state, keep capacity movable, and create environments that autoscalers can safely consolidate. 

Adi Fayer is a Senior Product Manager at Komodor

Hot Topics

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...

From Cleanup to Prevention: The Future of Cloud-Native Cost Optimization

Adi Fayer
Komodor

Cloud teams are under constant pressure to reduce spend without slowing development or increasing operational risk. They are deploying autoscalers, rightsizing workloads, enforcing resource requests, reviewing utilization dashboards, and building FinOps processes around cloud-native environments.

Yet the results often disappoint. The first wave of savings comes quickly, then progress stalls. Clusters still run larger than they should, while nodes remain partially empty but cannot be removed. Platform teams see idle capacity, but the infrastructure cannot safely release it. The finance department sees ongoing waste, but engineering teams insist the environment has already been optimized.

Both sides are right.

The next phase of savings will come from preventing inefficient cluster states before they form.

Why Reactive Optimization Hits a Ceiling

Most cost programs start after waste already exists. Rightsizing tools adjust workload requests. Autoscalers try to consolidate nodes. FinOps platforms show where spend is high. These capabilities are useful, but they operate on a cluster state that has already been shaped by earlier scheduling decisions.

That is the core limitation.

The Kubernetes scheduler places pods based on what fits at the moment. Autoscalers try to clean things up later by removing capacity that is no longer needed. But if workloads have already landed in ways that scatter unevictable pods, create fragmentation, or violate consolidation paths, the autoscaler inherits a problem it isn't designed to solve.

This explains why teams using Karpenter or Cluster Autoscaler may still see significant capacity remain idle or locked.

The Real Problem Is Cluster State

The problem with reactive optimization is that it often acts on waste that is already entrenched and has become difficult to remove.

By the time a dashboard shows unused CPU or memory, that capacity may already be scattered across nodes the cluster cannot safely drain. Some nodes may be pinned by configurations associated with Pod Disruption Budgets or affinity rules. On paper, the cluster has available capacity, but that capacity cannot be practically consolidated into fewer running machines.

This is why utilization improvements have a ceiling. Rightsizing, autoscaler tuning, and better dashboards can only act on the cluster state that already exists. They may expose unused capacity, but they cannot always undo the placement decisions, policies, and workload constraints that prevent nodes from being drained. When that happens, the waste remains structurally locked into the cluster.

Dashboards may show improvement. Teams may feel they have done the right work. But if inefficiently utilized nodes cannot be drained and terminated, infrastructure spend does not fall.

That's why the cost conversation has to move beyond how much capacity each workload requests and focus on whether the cluster can be consolidated using existing infrastructure to reduce idle resources.

This requires a different operating model.

Move Optimization Earlier

The most important shift is pushing cost intelligence closer to scheduling time.

Instead of waiting for the cluster to fragment and then trying to consolidate it later, platform teams need to influence placement before workloads land. That means evaluating not only whether a pod fits now, but whether placing it on a particular node will make the cluster harder to scale down later.

This is a proactive scaling mindset. It considers workload behavior, reliability constraints, eviction rules, autoscaler logic, and future drain scenarios before placement decisions create waste. If a node is likely to be removed, the platform should avoid deploying new workloads there. If certain pods are difficult to evict, they should be placed in dedicated nodes rather than scattered randomly. If complementary workloads can share capacity efficiently, the platform should account for that before idle resources become stranded.

The goal should be to keep the environment flexible enough for autoscalers to do their job more efficiently.

Detection and Prevention Must Work Together

Proactive placement is only part of the answer. Most enterprise clusters already contain years of accumulated constraints, exceptions, and workload patterns that block consolidation. Preventing new waste does not automatically eliminate old waste.

That means the modern cost model needs two loops.

The first loop detects existing blockers: workloads that prevent node removal, overly restrictive disruption policies, autoscaler settings that preserve waste, and nodes kept alive by a small number of unevictable pods. The second loop prevents new inefficiency by guiding placement decisions before those blockers spread.

When these loops work together, cost optimization becomes continuous state management. The platform is not merely identifying waste after the fact. It is keeping the cluster in a condition where savings can actually be realized.

This is a more mature view of cloud-native efficiency. It recognizes that the best cost decision is not always the cheapest immediate placement. It is the placement that preserves reliability while allowing infrastructure to scale down cleanly over time.

From Cleanup to Cost-Aware Design

The next layer of cloud-native savings will not come from pushing teams to rightsize harder. That work still matters, but it is not enough.

The larger opportunity is to engineer efficiency into the platform before waste takes hold. That requires treating scheduling, autoscaling, and reliability policy as part of the same cost system.

Cloud-native infrastructure was built to be dynamic, but many cost practices remain static and reactive. They focus on waste only after it appears, explain why spend is high after the bill arrives, then expect autoscalers to clean up placement decisions that were never optimized for future consolidation.

That model is reaching its limit. The next phase of cloud cost management needs to be proactive. It should prevent bad cluster state, keep capacity movable, and create environments that autoscalers can safely consolidate. 

Adi Fayer is a Senior Product Manager at Komodor

Hot Topics

The Latest

Rapid AI adoption and the unique ways AI workloads operate is redefining the scope and structure of what these teams must deliver. This shift is forcing organizations to rethink how they manage scale, automation, and control, according to The State of SRE and Platform Engineering 2026, a new report from Dynatrace ...

AI is usually talked about as a software tool, but it also depends heavily on the network behind it. Whether a company is using AI for chatbots, automation, monitoring, analytics, or employee support, all of that information has to move across the network in a reliable and secure way. That means AI is not just an application decision. It is also an infrastructure decision. Before organizations rush into AI, they should ask a simple question: Is our network ready to support it? ...

Enterprise AI often lacks governed access to where business processes actually execute. Without that access, AI agents may be able to reason, but they cannot operate reliably across enterprise workflows. For AI agents to effectively carry out workflows, they will require integration-layer context and controls. Organizations can implement these prerequisites by providing AI with managed access to the middleware layer ...

Enterprise networks rarely behave the same way for very long. A routing adjustment in one region may unexpectedly alter application performance in another. A cloud migration may introduce hidden dependencies that go unnoticed until an outage occurs. All the while, the network is managed by several different teams, each of whom use different tool sets — and as a result, have different views of the network ... There’s usually an engineer who remembers why traffic fails over a certain way between sites, or which transparent firewall was added where. The problem is that human memory cannot scale alongside enterprise-scale networks ...

Ask an infrastructure team how confident they are in their ability to govern AI, and most will tell you they've got it handled. A recent survey of 406 IT decision-makers and platform engineering leaders found 86% expressing exactly that confidence. Ask the same group whether they have a formal written AI governance policy, and the number drops to 30%, according to Spacelift's Infrastructure Automation Report ...

In MEAN TIME TO INSIGHT Episode 27, Shamus McGillicuddy, EMA VP of Research, Network Infrastructure and Operations, and Parker Hathcock, EMA Research Director covering IT Service/Operations (ServiceOps), discuss observability unification in modern IT operations ... 

Virtual Private Networks became a cornerstone of enterprise security at a time when corporate infrastructure looked very different from today ... For years, this model worked well. But the architecture behind VPNs assumed a centralized corporate environment—one where the network itself was the hub of activity. In a cloud — first world, that assumption no longer holds ...

Website outages get resolved just as fast in August as they do in November. I went looking for the opposite: the summer slowdown everyone assumes is there once the people who fix things are away. It isn't in the data we collected, covering 1.8 million confirmed outages across tens of thousands of websites ...

This year, many of the cloud infrastructure contracts signed in the early days of the AI boom will come up for renewal. As the year goes on, I anticipate we'll see a significant amount of cloud vendor swapouts and multi-cloud adoption, and the reason isn't just GPU depreciation. It's because they're tired of their current cloud providers ...

There's a moment the many observability teams have experienced days into bringing a new service into production: you realize that the vendor's claims of "intelligent" behavior included a large serving of hype. Their dashboards look nice until they don't, the failure modes are a black box, and no one on the team can confidently explain why the system did what it did at 2 am. Agentic AI is about to force every Ops team to relive that moment at web-scale until they start treating these systems as the dependencies they actually are ...