kubernetes

Autoscaling the Waste: Why Kubernetes Keeps Buying Capacity Your Applications Don’t Use

On this page

Key takeaways

  • Autoscaling controllers respond to resource requests, not application consumption. Oversized requests can cause nodes to be provisioned even when workloads use a fraction of reserved capacity.
  • Cluster Autoscaler and Karpenter both provision and consolidate based on request totals, not metrics server data. Low utilization does not automatically trigger node removal.
  • Nine categories of consolidation blockers, including Pod Disruption Budgets, affinity rules, topology spread constraints, and long termination grace periods, can prevent valid node removal even when nodes appear underutilized.
  • HPA calculates replicas using actual consumption divided by resource requests. Reducing requests increases reported utilization, which can trigger scale-out and offset efficiency gains.
  • The four layers that matter are: application consumption, resource requests, resource limits, and provisioned node capacity. Each layer has different mechanisms and different levers.
  • CPU throttling is a distinct signal from overprovisioning. A pod with low average utilization may still hit its CPU limit during bursts, and cutting requests in that case degrades performance rather than recovering waste. The lower utilization section covers this distinction in detail.
  • Coordinated workload and node optimization addresses both provisioning and consolidation. Changes at one layer alone produce incomplete results.

What Kubernetes autoscaling actually changes

Kubernetes provides autoscaling at three distinct layers. Each layer responds to different inputs, changes different things, and leaves other things unchanged. Understanding these distinctions is the starting point for any efficiency work.

Layer

Responds to

What it changes

What it does not change

Relationship to resource requests

HPA (Horizontal Pod Autoscaler)

CPU/memory utilization as % of requests, or external/absolute metrics

Replica count

Resource requests per pod; node count directly

Utilization = consumption / requests. Request size directly affects the utilization ratio and therefore HPA decisions when using percentage-based metrics.

VPA (Vertical Pod Autoscaler)

Observed historical resource usage per container

Resource requests and limits

Replica count; node count directly

VPA is the mechanism that adjusts requests. Changes from VPA feed into scheduler and node autoscaler decisions downstream. Kubernetes 1.33+ includes InPlacePodVerticalScaling (beta), which allows VPA to resize running pods without eviction or restart. Feature gate availability varies by managed Kubernetes provider; verify support on EKS, GKE, and AKS before relying on it.

Cluster Autoscaler / Karpenter

Pending pods (provisioning); request totals on nodes (consolidation)

Node count and node types

Resource requests per pod; replica count

Provisioning triggers when requests cannot be scheduled. Consolidation triggers when request totals on a node fall below a threshold and remaining pods can fit elsewhere.

None of these controllers directly measure whether your application is processing work at an appropriate rate. They respond to the configuration you provide. When requests are accurate and policies match demand, autoscaling keeps infrastructure well-aligned with workload needs. When requests are oversized or policies are misconfigured, autoscaling consistently maintains an inefficient state. Calling autoscaling “broken” in these cases is inaccurate. The system is simply doing what you configured it to do.

HPA and VPA anti-pattern: Running HPA and VPA simultaneously on the same CPU or memory metrics is a well-known anti-pattern that causes scaling thrash. HPA scales replicas up when utilization rises; VPA simultaneously changes requests, which shifts the utilization ratio that HPA reads, which then triggers another HPA cycle. The exact failure mode depends on VPA configuration. With controlledValues: RequestsOnly (VPA’s default), VPA lowers CPU requests but leaves limits unchanged – this causes CPU throttling as the container hits the unchanged limit during peak usage, not a replica oscillation loop. The oscillation pattern occurs when VPA is also adjusting limits proportionally, changing the utilization denominator HPA reads. To avoid this, use VPA in Off (recommendation-only) mode when HPA is scaling on CPU or memory. Alternatively, configure them on orthogonal signals: HPA on external or custom metrics, VPA adjusting requests based on observed consumption without active enforcement.

The efficiency gap behind the autoscaling debate

The gap between what applications consume and what infrastructure provisions is measurable. In the analyzed sample from Cast AI’s 2026 Kubernetes Optimization Report, average CPU utilization was 8%, down from 10% in the prior year’s data, and average memory utilization was 20%, down from 23% previously. These figures cover the January 1 to December 31, 2025 measurement period for CPU and memory. They do not represent all Kubernetes clusters globally.

Low utilization figures do not translate directly to an equivalent proportion of the infrastructure bill being eliminable. Workloads require headroom for traffic bursts. Latency-sensitive services need operating margin above steady-state consumption. Not all headroom is waste.

What the numbers reveal is that requests, policies, and scheduling constraints in many environments are not calibrated to actual demand. In the same analyzed sample, CPU overprovisioning jumped from 40% to 69% year over year. Memory overprovisioning reached 79%. These figures describe the gap between declared requests and observed consumption. The relevant question for each environment is: which portion of that gap is legitimate reserve, and which portion can be reduced without affecting application performance? One critical caveat before acting on those numbers: low average utilization and active CPU throttling require different responses. Throttling means a pod is hitting its CPU limit during bursts, not that its requests are too generous. The lower utilization section below covers this failure mode directly.

How oversized requests become additional nodes

Resource requests tell the Kubernetes scheduler how much capacity to reserve for a pod. The scheduler places pods on nodes based on these declared values, not on what pods are actually consuming at runtime. This creates a direct mechanical path from oversized requests to additional provisioned nodes.

Four distinct layers govern how capacity flows from application to infrastructure. First, application consumption is what the process actually uses at runtime. Second, resource requests specify the values in the pod spec that the scheduler uses to make placement decisions. Third, resource limits set the maximum amount of resources a container can consume before CPU is throttled or memory causes the container to be killed. Fourth, provisioned node capacity is the total allocatable resources across running nodes. Waste accumulates most often in the gap between layers one and two. VPA is the primary Kubernetes mechanism for closing that gap by adjusting requests toward observed consumption (see the VPA row in the table above). VPA in Recreate mode also interacts with PDBs: if a PDB prevents the eviction needed to apply new resource settings, VPA’s recommendation sits pending without being applied. Engineers may see a VPA recommendation that appears to not take effect – checking PDB constraints is the first diagnostic step.

Scheduling arithmetic

Consider a concrete example. A Deployment has 12 fixed replicas. Each replica requests 2 vCPU. Each node provides 8 vCPU allocatable to these workloads (allocatable capacity is typically 90-95% of raw node capacity after system-reserved and kube-reserved deductions). DaemonSet pods further reduce per-node allocatable capacity for standard workloads, since DaemonSets run on every node and their resource requests are deducted before workload scheduling. The scheduler therefore needs at least ceil(12 × 2 / 8) = 3 nodes to place all replicas.

Each replica actually consumes 0.3 vCPU at steady state. Total actual consumption is 3.6 vCPU across all replicas. But the scheduler does not look at consumption. It looks at requests. Three nodes remain provisioned because the declared requests require it.

Reducing each replica’s request to 1 vCPU creates a theoretical packing opportunity: ceil(12 × 1 / 8) = 2 nodes. This example assumes that memory remains nonbinding, the replica count stays fixed at 12, placement constraints allow consolidation, and other cluster overhead remains excluded. This is scheduling arithmetic. It identifies a potential packing outcome, not a guaranteed savings estimate. Actual node reduction depends on whether consolidation blockers are absent and whether the node autoscaler’s consolidation policy permits the drain.

Why the gap compounds at scale

The effect compounds across a large cluster. A cluster with dozens of Deployments carrying similar request inflation accumulates phantom capacity across many nodes. Both Cluster Autoscaler and Karpenter provision nodes to satisfy pending pods and consolidate nodes based on request totals. Neither tool reads the metrics server for provisioning decisions. The Kubernetes node autoscaling documentation makes this clear: provisioning is driven by scheduling feasibility, not utilization data.

This means request sizing is upstream of every autoscaling decision. Tuning HPA targets or consolidation policies on top of inflated requests produces marginal results at best. The more productive starting point is aligning requests with actual consumption before adjusting any autoscaling policy. The Datadog State of Cloud Costs report found that 83% of container costs in its analyzed sample went to idle resources, split between overprovisioned cluster infrastructure (54%) and oversized workload requests (29%). The request layer and the node layer each contribute independently to that gap.

Why low utilization does not automatically trigger scale down

Seeing low CPU utilization across a cluster does not mean node autoscaling will remove nodes. Several conditions must be true before Cluster Autoscaler or Karpenter marks a node as removable and begins draining it.

For Cluster Autoscaler, a node becomes a scale-down candidate only when its pods can be rescheduled on other nodes without violating any scheduling constraint. The controller then waits for the default scale-down-unneeded-time of 10 minutes before draining. Karpenter’s consolidation behavior is configurable per NodePool and may require explicit enablement depending on version and configuration.

Nine categories of conditions commonly block consolidation:

  • Minimum node counts or capacity floors. NodePool or node group minimums keep nodes running regardless of utilization metrics.
  • Pod Disruption Budgets (PDBs). A PDB with a tight minAvailable or maxUnavailable setting can make every node ineligible for draining. If evicting any pod would drop replicas below the PDB threshold, the node stays.
  • Affinity and anti-affinity rules. Anti-affinity rules that spread replicas across nodes can block consolidation. If no other node satisfies the scheduling constraint, the pod cannot move.
  • Topology spread constraints. Constraints enforcing zone or hostname distribution can pin pods to specific nodes, preventing their removal.
  • Storage and placement requirements. Pods with local PersistentVolumes or node-selector constraints may not be movable to other nodes in the cluster.
  • Resource fragmentation. A node can have low CPU utilization while being fully utilized for memory. If other nodes cannot accommodate the memory required by its pods, the node cannot be consolidated, even when its overall CPU utilization is low.
  • Controller-specific consolidation settings. Both Cluster Autoscaler and Karpenter expose per-controller configuration that can disable or delay consolidation independently of utilization state.
  • Remaining pods that cannot fit elsewhere. If any pod on a node cannot be placed on another node given current cluster state, Cluster Autoscaler will not drain that node regardless of utilization.
  • Long terminationGracePeriodSeconds. Pods that drain persistent connections before shutdown commonly carry grace periods of 5 to 10 minutes. These long termination windows can prevent a node drain from completing within the consolidation timeout. Cluster Autoscaler and Karpenter both impose limits on how long they wait for a drain to finish. A node with slow-terminating pods may be abandoned mid-consolidation and remain provisioned. Check terminationGracePeriodSeconds on any workloads that handle long-lived connections before expecting consolidation to complete within a normal window.

When several of these conditions apply simultaneously, nodes persist even when aggregate utilization appears low. Diagnosing why consolidation is not happening requires inspecting each blocker category, not just examining utilization dashboards.

Changing requests can also change HPA behavior

The Kubernetes HPA documentation specifies that CPU utilization is computed as actual consumption divided by resource requests, expressed as a percentage. The formula for replica adjustment is:

desiredReplicas = ceil(currentReplicas x currentMetricValue / desiredMetricValue)

The HPA controller runs on a 15-second sync interval. This formula applies specifically when HPA is configured to use CPU or memory utilization as a percentage of resource requests. HPA can also use absolute metric values or external metrics sources. In those cases, resource request size has no effect on the utilization calculation, and this coupling does not apply.

For utilization-based HPA, reducing requests has a direct mechanical consequence. If a pod consumes 0.4 vCPU and the request is 2 vCPU, reported utilization is 20%. Reduce the request to 0.5 vCPU without changing actual consumption, and reported utilization becomes 80%. If the HPA target is 70%, the controller now sees a workload above threshold and scales out additional replicas.

That scale-out increases pod count. More pods require more scheduling capacity. The efficiency gain from reducing requests can be partially or fully offset by HPA-driven replica increases. This is not a flaw in HPA. It is an intended property: the controller keeps utilization near the configured target. Request adjustments and HPA targets must therefore be evaluated and coordinated together, not applied independently.

Before adjusting requests on workloads with active utilization-based HPAs, verify the HPA target and calculate the expected utilization ratio after the change. Apply changes to a subset of workloads first, then monitor replica counts, HPA events, and application performance metrics before proceeding further.

A practical audit of autoscaling waste

Diagnosing autoscaling waste requires collecting evidence at each layer before drawing conclusions. The following six steps provide a structured approach. Work through them in order; each step informs the next.

Step 1: Establish the baseline

Before changing anything, record the current state. Capture requested and consumed CPU and memory per pod and namespace, provisioned node count and type, total allocatable capacity, and current cost per relevant unit of work. Units of work vary by service: per HTTP request handled, per batch job completed, or per hour by service tier. A baseline without cost per unit of work cannot confirm whether later changes produced real savings or just shifted metrics.

Step 2: Compare requests with consumption

Pull requested and consumed CPU and memory per pod using kubectl top pods and your observability platform. Look for pods with sustained large gaps between requested and consumed resources. Flag pods with CPU throttling or OOM events separately: these pods are hitting resource limits during bursts and require investigation before any request reduction. A pod hitting its CPU limit consistently is not a candidate for rightsizing down. Also consider the risk if you plan to use VPA Recreate mode to apply tighter limits. VPA Recreate evicts and restarts pods with updated resource values. If the new limits do not account for startup memory peaks, OOMKill can follow the eviction immediately. Analyze startup memory behavior before tightening limits on any pod with Recreate mode.

Step 3: Trace scale-up decisions

Examine node provisioning events in cluster event logs. Run kubectl get events -n <namespace> –field-selector reason=TriggeredScaleUp to identify which pending pods triggered each scale-up event and what their resource requests were. Run this per namespace or add –all-namespaces to see cluster-wide scale-up events. Cluster-wide output can be noisy in large clusters. Then check whether those pods actually ran near their requested resources after placement. A scale-up triggered by a pod that subsequently ran at 5% of its requested CPU is a signal that the request was the constraint, not actual demand.

Step 4: Inspect consolidation blockers

For each node that has remained underutilized beyond the scale-down-unneeded-time threshold, work through the nine blocker categories systematically. Start with kubectl get pdb -A to list all Pod Disruption Budgets and identify any with thresholds that could block eviction. Use kubectl describe node <node> to inspect conditions, allocated resource totals, and which pods are pinned to the node. Also check affinity rules, topology spread constraints, node group minimums, and storage requirements. Examine resource fragmentation: a node may be low on CPU but fully subscribed in memory. Finally, check terminationGracePeriodSeconds on workloads with long-lived connections, as slow-terminating pods can cause drain operations to time out before consolidation completes.

Step 5: Pilot coordinated changes

Apply request adjustments to a subset of workloads. For each adjusted workload with a utilization-based HPA, monitor replica counts and HPA events alongside utilization metrics. Record pending pod duration before and after the change. Track node provisioning and removal events over a period that includes representative traffic, including peak hours and any scheduled batch workloads. Do not apply changes cluster-wide until pilot results are clear.

Step 6: Verify the outcome

After changes, confirm four things. First, CPU throttling and OOM events have not increased. Second, application latency, errors, and throughput remain within acceptable bounds. Third, node count or provisioned capacity has measurably changed. Fourth, cost per relevant unit of work has improved.

If application metrics worsen, roll back before investigating further. To revert resource request changes, restore the previous values directly in the deployment spec:

kubectl set resources deployment/<name> –requests=cpu=2000m,memory=512Mi

Note: this is a break-glass manual override that updates the live spec directly. In GitOps environments, the authoritative fix is patching the Helm values file or kustomize overlay, then allowing the reconciler to sync. The kubectl set resources command will be overwritten on the next reconcile unless the source of truth is also updated.

For changes applied via GitOps or versioned manifests, redeploy the prior spec. Do not wait for issues to resolve themselves after a request reduction.

If cost per unit of work did not change despite lower utilization numbers, check whether the cluster runs on committed spend. Savings Plans, Reserved Instances, and Enterprise Discount Programs lock in charges for a defined period regardless of actual capacity consumed. Scaling down nodes during an active commitment does not reduce the current bill. The savings only appear in billing after the commitment period ends or when on-demand capacity is the marginal resource being removed. This is the most common reason a measurable reduction in provisioned capacity does not appear in the next invoice.

Diagnostic reference: three common symptoms

Symptom

Possible explanation (hypothesis)

Evidence to inspect

Low CPU usage but no node removal

Consolidation blockers are preventing node drain; or per-node request totals are not low enough to trigger scale-down thresholds

PDB settings, affinity and topology constraints, node group minimums, scale-down-unneeded-time setting, request totals per node compared to allocatable capacity

Pending pods despite spare aggregate capacity

Aggregate capacity is available but not where the scheduler needs it; pending pods require specific node types, zones, or memory ratios that are locally exhausted

Pending pod events and scheduling failure reason, node selectors and affinity rules on pending pods, available allocatable capacity broken down by node type and zone

More nodes after a deployment update

New deployment version carries higher CPU or memory requests than the previous version; or replica count increased; or updated anti-affinity rules now prevent existing packing

Diff the deployment manifest for request changes between versions, compare new and old replica counts, check anti-affinity rules in the updated spec

These are hypotheses to investigate, not automatic diagnoses. Multiple explanations can apply to the same symptom simultaneously. The evidence column points toward where to distinguish between a benign cause and a correctable one.

What coordinated workload and infrastructure optimization should deliver

Addressing autoscaling waste at the workload layer and the infrastructure layer independently produces incomplete results. Request adjustments without consolidation changes leave nodes provisioned. Node policy changes without request adjustments may trigger unwanted HPA behavior. Both layers need to move together.

At the workload layer, coordinated optimization means rightsizing resource requests based on observed consumption patterns, with appropriate headroom for burst, startup, and peak behavior. At the infrastructure layer, it means reviewing consolidation settings and resolving blockers that prevent valid node removal.

Cast AI’s workload optimization and node autoscaling components operate across both layers. The workload optimization component analyzes consumption history and recommends or applies request adjustments. The node management component handles provisioning and consolidation based on updated request totals. Because both components share a consistent view of workload state, request adjustments feed directly into updated consolidation targets without requiring separate manual coordination.

The workload optimization component measures p95 CPU consumption rather than averages, which means burst behavior is accounted for before recommendations are generated. For HPA-managed workloads scaling on CPU utilization, Cast AI avoids changing CPU requests directly, since doing so would shift the utilization ratio HPA reads and potentially trigger unintended scale-out. Before any automated changes are applied, recommendations are available in a read-only mode: the system surfaces what it would change and why, based on observed consumption patterns, so teams can audit recommendations against application behavior before enabling automation.

The expected output of coordinated optimization is not a utilization percentage target. It is appropriate capacity for each workload at acceptable application performance, with unnecessary headroom removed and consolidation blockers addressed systematically. The measurable signal is cost per unit of work, not average cluster utilization.

When low utilization masks a throttling problem

A specific scenario requires care before adjusting requests: when an application is already running close to its performance boundary, existing headroom is not waste. It is operating margin.

CPU throttling metrics reveal this condition directly. Consider a pod consuming 0.3 vCPU with a request of 0.5 vCPU and a limit of 0.5 vCPU. Average utilization looks low. But the pod is already hitting its CPU limit during processing cycles. Reducing the request to match average consumption will not reduce throttling. Reducing the limit below actual peak demand will worsen it and may degrade latency or throughput.

Latency-sensitive services, services that batch-process large datasets, and services handling irregular high-volume spikes all require careful analysis before any request changes. A low average utilization figure can coexist with frequent, legitimate bursts that the average obscures. P95 or P99 consumption data is more informative than averages for these workloads, because averages smooth over the peaks that actually stress the resource boundary.

The audit step on verifying outcomes exists precisely for this reason. If application latency, errors, or throughput worsen after adjustments, roll back immediately using the procedure in Step 6. Appropriate capacity at acceptable performance is the objective. Infrastructure efficiency does not override application correctness.

Conclusion

Kubernetes autoscaling does not cause resource waste. It maintains whatever configuration you provide. When resource requests, scheduling constraints, and controller policies do not reflect actual demand, autoscaling faithfully preserves the resulting inefficiency. Tracing the gap requires examining all four layers: application consumption, resource requests, scheduling behavior, and node consolidation. The six-step audit above provides a structured starting point. For teams ready to move from diagnosis to action, Cast AI’s Kubernetes cost optimization documentation covers coordinated implementation options in detail.

Frequently Asked Questions

Why can Kubernetes autoscaling leave a cluster overprovisioned?

Autoscaling controllers respond to resource requests, scheduling constraints, and policy settings rather than to actual application consumption. When resource requests are set higher than workloads consume, the scheduler reserves more capacity than needed. Node autoscalers provision nodes to satisfy those requests. Unless consolidation conditions are met and blockers are absent, that excess capacity persists. The result is a cluster that autoscales correctly by its own rules while remaining inefficient relative to actual demand.

Does Cluster Autoscaler use actual CPU consumption to size nodes?

No. Cluster Autoscaler provisions nodes when pending pods cannot be scheduled given current resource requests. It consolidates nodes when request totals on a node fall below a threshold and remaining pods can fit on other nodes. Neither provisioning nor consolidation uses the metrics server or actual CPU consumption data. Karpenter follows the same model. Autoscaling alone will not scale down a node whose pods collectively request 6 vCPU but consume only 0.5 vCPU unless the request totals meet the consolidation threshold and no blockers prevent removal.

What is the difference between rightsizing and node autoscaling?

Rightsizing adjusts the resource requests and limits declared on individual pods or containers to better match observed consumption. Node autoscaling provisions or removes nodes based on scheduling feasibility and node-level request totals. The two mechanisms are complementary but operate at different layers. Rightsizing changes the inputs that node autoscaling responds to. Without rightsizing, node autoscaling operates on whatever requests are declared, which may significantly overstate actual needs. Without node autoscaling, rightsizing alone does not remove already-provisioned infrastructure.

Why are underutilized nodes not being removed?

Several conditions can block node removal. Pod Disruption Budgets may prevent pod eviction if removing a pod would drop replicas below a defined minimum. Affinity or anti-affinity rules may make remaining pods unschedulable on other nodes. Topology spread constraints may pin pods to specific zones or hostnames. Node group minimums may set a floor that keeps nodes running regardless of utilization. Resource fragmentation, where a node is low in CPU but high in memory, can also prevent pod migration. Long terminationGracePeriodSeconds values are another common blocker: pods configured to drain persistent connections over 5 to 10 minutes can cause drain operations to time out before consolidation completes, leaving the node provisioned. Check each of these conditions systematically before concluding a node should be removable. If VPA runs in Recreate mode, PDB constraints can block the pod eviction required to apply updated resource settings. The VPA recommendation may appear to have no effect until you resolve the PDB conflict.

Can changing CPU requests affect HPA scaling?

Yes, when HPA is configured to use CPU utilization as a percentage of resource requests. HPA computes utilization as actual consumption divided by the declared request. If requests decrease and actual consumption stays constant, reported utilization increases. A higher utilization percentage can trigger HPA to scale out additional replicas, which increases pod count and may require more nodes. This coupling means request changes and HPA targets need to be evaluated together. If HPA is using absolute metrics or external metrics rather than utilization percentages, changing requests has no effect on HPA behavior.

How do you verify that autoscaling changes reduced costs safely?

Verification requires confirming changes on both the application side and the cost side. On the application side, check that CPU throttling and OOM events have not increased, and that latency, error rates, and throughput remain within acceptable bounds over a representative traffic period including peak hours. On the cost side, confirm that node count or provisioned capacity has measurably decreased, and that cost per relevant unit of work has improved. A reduction in utilization percentages alone does not confirm savings. If the cluster runs on committed spend such as Savings Plans, Reserved Instances, or Enterprise Discount Programs, scaling down nodes may not reduce the bill until the commitment period ends. This is the most common reason a measurable drop in provisioned capacity does not appear in the next invoice. Actual billing changes depend on capacity removal affecting on-demand charges, instance type changes, or similar events that affect what you are charged for.

Leave a Reply

Your email address will not be published. Required fields are marked *

Chat on WhatsApp