GKE 1.37 Scales to Zero—but CPU Metrics Cannot Wake It Up

|Author: QUASA Editorial Team|5 min read| 1
GKE 1.37 Scales to Zero—but CPU Metrics Cannot Wake It Up

On September 23, 2026, Google introduced native HPA scale-to-zero in GKE 1.37. An eligible Deployment can shed its last idle Pod and regain replicas when demand returns. The release pairs that change with shared capacity buffers intended to reduce the wait for compute when a workload restarts.

An independent account of the GKE rollout identifies intermittent batch processors and event-driven workers among the likely uses. Their suitability turns on a specific question: can the autoscaler measure incoming work while no worker Pods exist? A queue can retain messages during that interval; CPU usage cannot reveal demand until a Pod is running.

The wake-up metric must exist without Pods

The GKE configuration guide requires a control plane and nodes running version 1.37 or later, an autoscaling/v2 HPA with minReplicas: 0, and at least one External or Object metric. CPU and memory are Resource metrics gathered from running Pods. At zero replicas, those readings cannot, by themselves, trigger a return to service.

The Pub/Sub example uses undelivered messages in a subscription as the demand signal. An AutoscalingMetric resource exposes that measure to the HPA, which targets a worker Deployment. When the queue empties, the HPA can remove the workers; when messages arrive, the subscription still holds a measurable backlog that can prompt the HPA to restore them.

An Object metric can serve the same purpose if its value remains available when the target workload stops and changes as demand returns. A CPU-only configuration fails that test: no running worker means no worker CPU reading, even if jobs are waiting elsewhere. Setting the minimum replica count to zero does not create an independent signal.

Version and namespace checks decide whether the configuration qualifies

The release requirement covers both newly created clusters and existing clusters being upgraded. Updating the control plane alone is insufficient when its nodes still run an earlier version. The HPA must also use the specified API version and an eligible metric type; changing only its minimum replica field leaves those prerequisites unmet.

The AutoscalingMetric, HorizontalPodAutoscaler and target Deployment must be in the same Kubernetes namespace. They are separate resources with separate roles: the metric resource defines what GKE reads, the HPA sets the scaling rule, and the Deployment supplies the Pods. In the Pub/Sub configuration, the HPA references the external metric name exposed by AutoscalingMetric and identifies the worker Deployment as its scale target.

This creates a practical go-or-no-go test for replacing an additional autoscaling component. If the workload’s demand is available as a usable External or Object metric through GKE’s managed metric path, native HPA can control the move to and from zero without a separate KEDA operator or third-party metrics adapter. If the required demand signal is unavailable through that path, removing the existing component would leave the workload without a reliable trigger to restart.

An HPA-controlled stop must remain active

A Deployment showing zero replicas has not necessarily entered a recoverable idle state. For an HPA-initiated reduction, the relevant conditions are ScaledToZero: True and ScalingActive: True. The first records that the HPA made the move to zero; the second indicates that it can still calculate a replica count from the configured metric. The conditions are visible with kubectl describe hpa async-worker-hpa in the Pub/Sub example.

Manually scaling the Deployment to zero produces a different state. The HPA pauses to avoid conflicting changes, and ScalingActive becomes False; scaling the Deployment back to at least one replica resumes autoscaling. An operator who intends the HPA to wake a worker should therefore let the HPA make the final scale-down rather than set the Deployment to zero independently.

The external metric must also remain readable during the idle period. If it becomes missing or invalid, the HPA cannot calculate the replica count needed when demand returns. The queue may still preserve work, but preserving work and starting capacity are separate functions.

Capacity buffers address the next wait

A rising demand metric can cause the HPA to request a Pod, but the Pod still needs compute capacity, a container image and application startup time before it can process work. Without suitable capacity already available, provisioning a node adds another delay. GKE capacity buffers retain shared warm compute that a returning Pod can claim, addressing the infrastructure part of that cold start.

An active buffer makes capacity available to workloads returning from zero; a standby buffer can replenish the active pool as demand continues. Sharing that pool can spare each intermittent workload from keeping its own idle replica solely to reserve compute. The buffer itself retains infrastructure resources, so its size affects the balance between idle capacity and the time a returning Pod waits for a node.

For a queue worker, the decision depends on both parts of the recovery path: the metric must reveal pending work without Pods, and capacity must become available soon enough for the workload’s response requirement. Native HPA can remove an autoscaling component when its managed metric path covers that workload. It cannot eliminate application startup time, and workloads with tighter response requirements may still need warm capacity when the queue begins to fill.

Also read:

Share:

Subscribe to our newsletter

Get the latest Web3, AI, and crypto news delivered straight to your inbox.

0