Kubernetes autoscaling adjusts workload replicas or worker-node capacity as demand changes. The two layers solve different problems:
HPA changes the number of Pods. Node autoscaling changes the amount of worker capacity available to schedule those Pods.
For small teams, the safe operating model is:
measure → set realistic requests → scale Pods → observe Pending Pods → scale workers → consolidate carefully
Raff Technologies managed Kubernetes supports worker-pool autoscaling between the minimum and maximum values configured per pool. The application team still owns HPA signals, replica bounds, workload requests, readiness behavior, and the cost ceiling.
This guide owns Kubernetes autoscaling strategy: Horizontal Pod Autoscaler, node autoscaling, scale-to-zero, scheduling interaction, scale-down stability, and cost boundaries.
What is Kubernetes autoscaling?
Kubernetes autoscaling changes resources automatically in response to workload demand.
The three common layers are:
| Autoscaling layer | What changes | Typical signal |
|---|---|---|
| Horizontal Pod Autoscaler (HPA) | Number of Pod replicas | CPU, memory, custom, object, or external metrics |
| Vertical Pod Autoscaler (VPA) | Pod resource requests/recommendations | Historical and current resource use |
| Node autoscaling | Number of worker nodes | Unschedulable Pods and removable excess capacity |
For most small teams, the key production combination is HPA + node autoscaling.
HPA answers:
How many replicas should this workload run?
Node autoscaling answers:
Does the cluster have enough compatible worker capacity to place them?
These decisions are connected, but they are not interchangeable.
How HPA works
The HorizontalPodAutoscaler periodically reads metrics and adjusts the desired replica count of a scalable workload such as a Deployment or StatefulSet.
With CPU utilization, the important relationship is approximately:
current utilization = measured CPU usage / requested CPU
That means resource requests directly affect the signal.
Example:
Measured CPU = 200m CPU request = 100m Utilization ≈ 200%
But:
Measured CPU = 200m CPU request = 500m Utilization ≈ 40%
The application is doing the same amount of CPU work. Only the request changed.
This is why Kubernetes Requests vs Limits must be correct before aggressive HPA tuning.
Bad requests create bad autoscaling signals.
HPA replica calculation is a ratio problem
At a high level, HPA calculates desired replicas from the relationship between current and target metrics.
A simplified form is:
desired replicas ≈ current replicas × current metric / target metric
Example:
current replicas = 4 current CPU utilization = 80% target CPU utilization = 50% desired replicas ≈ 4 × 80 / 50 ≈ 6.4
The controller then applies Kubernetes autoscaling behavior, tolerance, missing-data handling, stabilization, and configured min/max bounds.
Do not use this simplified formula as a billing prediction. It is useful for understanding why the controller wants more or fewer replicas.
Requests must exist for resource-utilization HPA
For CPU or memory utilization targets, HPA needs meaningful resource requests.
If a container does not have the relevant request, Kubernetes may be unable to calculate utilization for that Pod correctly.
Before enabling CPU-based HPA, verify:
kubectl get deployment <name> -o yaml kubectl get hpa kubectl describe hpa <name>
Check:
- requests are present;
- metrics are available;
- current and target values make sense;
- minReplicas and maxReplicas are intentional;
- HPA conditions do not show metric errors.
If HPA does not scale, inspect the metric path before increasing replica limits.
Use the metric that represents demand
CPU is convenient, but it is not always the right signal.
Different workloads may scale better from:
- CPU utilization;
- memory;
- request rate;
- queue depth;
- concurrent jobs;
- custom application metrics;
- external metrics.
Examples:
| Workload | Useful signal to evaluate |
|---|---|
| CPU-bound API | CPU |
| queue worker | queue depth |
| request-driven web tier | request rate or CPU |
| background processor | backlog / external metric |
| memory-bound service | memory, with caution |
| event consumer | lag or queue depth |
A metric should represent work that more replicas can actually help process.
If a database bottleneck causes API latency, adding more API Pods may make the database problem worse.
Autoscale the layer that owns the constraint.
HPA scale-up and scale-down should not behave identically
Scale-up often needs to react faster than scale-down.
Removing replicas too quickly can create oscillation:
demand falls → replicas removed → traffic rises slightly → replicas recreated → demand falls → replicas removed again
Kubernetes HPA includes behavior controls and a default downscale stabilization window. In current Kubernetes, the default downscale stabilization window is five minutes unless configured otherwise.
The autoscaling/v2 API allows teams to configure scale-up and scale-down behavior, including rate policies and stabilization.
For a small team, a good starting principle is:
scale up fast enough to protect service quality; scale down slowly enough to prove the capacity is really excess.
Tune from production behavior rather than copying one universal policy.
Kubernetes 1.37 can scale suitable HPA workloads to zero
Kubernetes 1.37 added Beta support for HPA scale-to-zero and enables it by default.
This is useful for workloads such as:
- queue consumers;
- event-driven workers;
- batch-style processors;
- workloads whose demand is represented by an object or external metric.
A suitable HPA can reduce a workload to zero replicas and later scale it back when the metric changes.
This can remove the last idle Pod for workloads that do not need permanent warm capacity.
But there is an important trade-off:
scale-to-zero introduces cold-start latency.
Kubernetes Services do not buffer requests while no Pods are ready. An ordinary HTTP application generally needs another component to queue or buffer demand if zero replicas are allowed.
Use scale-to-zero where work can wait safely, not simply because zero idle cost sounds attractive.
HPA tolerance reduces small scale oscillations
Kubernetes 1.37 also stabilizes configurable HPA tolerance.
Tolerance defines a band around the target metric where small variations do not trigger scaling.
Conceptually:
target CPU = 60% small movement around 60% → no scale action meaningful movement beyond tolerance → autoscaler considers a replica change
This helps prevent tiny metric fluctuations from constantly changing replica counts.
The correct tolerance depends on workload variability and how costly replica churn is.
Node autoscaling starts with unschedulable Pods
Node autoscaling operates below HPA.
A common scale-up sequence is:
traffic rises → HPA requests more replicas → scheduler tries to place new Pods → existing workers cannot fit them → Pods become Pending → node autoscaler evaluates compatible capacity → new worker joins → Pending Pods schedule
The important word is compatible.
A Pending Pod can be unschedulable because of:
- insufficient CPU;
- insufficient memory;
- node affinity;
- nodeSelector;
- taints and tolerations;
- topology rules;
- storage constraints;
- node-pool maximums;
- incompatible worker shapes;
- unavailable provider capacity.
Adding any random node does not necessarily solve the problem.
Use:
kubectl describe pod <pending-pod>
Read the scheduler events before assuming a Pending Pod means "add more nodes."
HPA and node autoscaling form one feedback loop
Consider:
- 3 normal replicas;
- HPA min = 3;
- HPA max = 12;
- one worker pool with autoscaling;
- enough existing headroom for 5 replicas.
Traffic increases.
HPA may request 6 replicas.
Five fit immediately.
The sixth stays Pending.
Node autoscaling then decides whether compatible capacity should be added.
This distinction matters for cost:
more Pods do not automatically mean more nodes.
If HPA growth fits inside existing headroom, paid worker capacity may not change.
Node cost increases when new worker capacity is actually required.
The reverse path is:
demand falls → HPA reduces replicas → scheduler has fewer Pods to place → nodes become underused → node autoscaler evaluates consolidation → removable workers leave
Scale-down is deliberately more conservative because workloads may need to move and disruption rules must be respected.
Node-pool architecture constrains autoscaling
Node autoscaling can only add useful capacity if the new node is eligible for the Pending workload.
For example:
Memory workload → required node affinity: pool=memory → memory pool at max nodes → general pool has spare capacity
The workload can still remain Pending.
Total cluster capacity is irrelevant if scheduling rules prevent the Pod from using it.
This is why Kubernetes Node Pools and autoscaling should be designed together.
For every pool, define:
- worker shape;
- minimum nodes;
- maximum nodes;
- workload eligibility;
- dedicated taints if required;
- expected failure reserve;
- scaling reason.
Too many pools can make autoscaling less efficient by fragmenting capacity.
The safe minimum fleet comes before autoscaling
Autoscaling does not create capacity instantly.
Scale-up requires:
- unschedulable demand to exist;
- the autoscaler to evaluate it;
- infrastructure provisioning;
- the node to join;
- system Pods to start;
- the node to become Ready;
- application Pods to schedule.
Your minimum worker count should therefore support the workload conditions that must survive immediately.
That may include:
- baseline traffic;
- one-worker failure;
- normal deployment surge;
- maintenance drain;
- always-on replicas.
Autoscaling should handle additional demand above the safe baseline.
Use Kubernetes Cluster Sizing to calculate that baseline.
Cost control comes from explicit bounds
Autoscaling does not inherently reduce cost.
It automates capacity inside the rules you provide.
Important boundaries include:
| Control | What it limits |
|---|---|
| HPA minReplicas | Minimum running application capacity |
| HPA maxReplicas | Maximum workload expansion |
| Worker minimum | Paid infrastructure floor |
| Worker maximum | Maximum node expansion |
| Requests | Scheduler reservation and utilization math |
| Node shape | CPU/memory added per scale event |
| Scale-down behavior | How quickly capacity is removed |
Every minimum and maximum should have a reason.
Examples:
HPA minimum = 3
Reason: application must preserve three warm replicas.
Worker minimum = 2
Reason: baseline workloads and maintenance requirements need two workers.
Worker maximum = 8
Reason: expected peak demand fits inside eight nodes and this creates an explicit infrastructure-cost ceiling.
If nobody can explain why a maximum is 50 rather than 10, the autoscaling policy is incomplete.
Cost increases often begin with inaccurate requests
Suppose a workload uses 300m CPU per Pod but requests 1 CPU.
The scheduler reserves capacity as if each Pod needs 1 CPU.
HPA adds more replicas.
The existing workers fill earlier.
Pods become Pending earlier.
Node autoscaling adds another worker earlier.
The autoscaler is behaving correctly.
The bad input is the request.
The reverse problem also matters. Requests that are far too low can pack too many Pods onto workers and create contention.
Use production observations to compare:
- actual usage;
- requests;
- limits;
- replicas;
- worker utilization;
- Pending Pod events;
- node scale events.
Rightsizing and autoscaling should be tuned together.
How Raff Kubernetes worker autoscaling works
Raff Kubernetes currently supports worker-pool autoscaling between the minimum and maximum values configured per pool.
The public product behavior is:
- scale up in response to Pending Pods that need capacity;
- scale down after sustained low load;
- set min and max independently per worker pool;
- scale manually when desired.
The current platform also supports multiple node pools, each with its own plan and node count.
Current shared-CPU workers include:
| Worker | vCPU | RAM | Monthly price |
|---|---|---|---|
| Starter | 1 | 2 GB | $9.99 |
| Standard | 2 | 4 GB | $17.99 |
| Performance | 4 | 8 GB | $33.99 |
| High Memory | 8 | 16 GB | $71.99 |
| Large | 8 | 32 GB | $119.99 |
| Scale | 16 | 64 GB | $229.99 |
Raff also offers CPU-Optimized Kubernetes workers for latency-sensitive workloads.
The standard control plane is currently $0/month, while three-master HA is $30/month.
For current worker pricing and product behavior, use the Raff Kubernetes product page.
The practical relationship is:
HPA / workload replicas ↓ resource requests ↓ scheduler fit ↓ Pending Pods ↓ worker-pool autoscaling ↓ paid worker capacity
That makes requests, replica limits, and worker-pool maximums direct cost controls.
Autoscaling troubleshooting matrix
When autoscaling does not behave as expected, identify the layer first.
| Symptom | Check first |
|---|---|
| CPU rises but HPA stays unchanged | HPA metrics, requests, conditions, maxReplicas |
| HPA adds replicas but Pods stay Pending | scheduler events |
| Pending Pod says insufficient memory | worker capacity / requests |
| Pending Pod says affinity mismatch | node-pool placement |
| Pool is at maximum | max-node boundary |
| Nodes scale but Pods still cannot run | taints, affinity, storage, topology |
| Workers do not scale down | minimum nodes, utilization, disruption, rescheduling feasibility |
| Replicas oscillate | HPA tolerance, stabilization, startup behavior |
| Cost rises without traffic growth | requests, replica floor, node minimums, fragmentation |
Do not debug HPA, scheduling, and infrastructure scaling as one black box.
Autoscaling checklist
Before enabling production autoscaling:
HPA
- CPU/memory requests are meaningful.
- The selected metric represents useful demand.
- minReplicas protects required warm capacity.
- maxReplicas is intentional.
- readiness/startup behavior is correct.
- scale-down behavior is conservative enough.
Worker autoscaling
- The safe minimum fleet is defined.
- Each pool has explicit min/max node counts.
- Worker shape matches the workload.
- Scheduling constraints are understood.
- Pool maximums fit expected peak demand.
- Failure capacity is not dependent on instant scale-up.
Cost
- Requests are reviewed against actual usage.
- Node-pool fragmentation is limited.
- Maximum worker count creates a known cost ceiling.
- Scale events can be correlated with workload demand.
Final recommendation
Treat Kubernetes autoscaling as one coordinated control loop.
Use HPA to change replica demand. Use node autoscaling to add compatible worker capacity when those replicas cannot be scheduled. Keep realistic resource requests between the two layers, define explicit minimums and maximums, and make scale-down more conservative than scale-up.
For workloads that can tolerate cold starts and use suitable external or object metrics, Kubernetes 1.37 HPA scale-to-zero can eliminate the last idle replica. For ordinary request-driven services, preserve warm capacity unless another layer safely buffers demand.
Continue with Kubernetes Requests vs Limits, Kubernetes Cluster Sizing, Kubernetes Node Pools, and Kubernetes Cost Optimization to tune the inputs and boundaries around autoscaling.