Kubernetes cluster sizing is the process of choosing enough allocatable worker capacity for normal workload demand, deployments, maintenance, failures, and expected growth.
The safest small-team model is:
required worker capacity = workload requests + system overhead + rollout overlap + failure reserve + planned growth
Do not size a cluster from the raw CPU and memory printed on a node plan alone. Kubernetes schedules Pods against Node Allocatable, and workload requests, placement constraints, DaemonSets, and failure headroom determine how much capacity is actually usable.
Raff Technologies separates the managed control plane from worker capacity, so teams can size workers from application demand rather than treating control-plane resources as workload capacity.
This guide owns the cluster-sizing decision: requests, allocatable capacity, node count, worker shape, N+1 failure tolerance, rollout headroom, and autoscaling boundaries.
Start with Node Allocatable, not raw node capacity
A node's advertised capacity is not the same as the resources available to application Pods.
Kubernetes exposes Node Allocatable after accounting for resources reserved for the operating system, Kubernetes system daemons, and eviction thresholds.
The scheduler uses allocatable capacity as the ceiling for Pod placement.
The sizing chain is:
raw node capacity * system-reserved resources * kube-reserved resources * eviction reserve = Node Allocatable Node Allocatable * DaemonSet / platform requests * application requests = remaining schedulable capacity
This is why two nodes with the same advertised vCPU and memory can provide different effective application capacity if their system workloads and reservations differ.
For an existing cluster, inspect the real value rather than estimating it:
kubectl get nodes kubectl describe node <node-name>
Look for the Capacity and Allocatable sections.
Use the measured allocatable values in capacity planning.
Resource requests are the scheduler's primary sizing signal
Kubernetes does not place Pods from observed average CPU alone.
The scheduler primarily considers resource requests when deciding whether a Pod fits on a node.
For every production workload, collect:
- CPU request per Pod;
- memory request per Pod;
- minimum replica count;
- maximum or expected rollout replica count;
- node-selector, affinity, taint, and topology constraints;
- required persistent-storage topology;
- DaemonSet overhead on eligible nodes.
Then calculate baseline requested capacity.
For one workload:
baseline CPU request = CPU request per Pod × minimum replicas baseline memory request = memory request per Pod × minimum replicas
Repeat for all workloads that must run at the same time.
Observed usage still matters because requests should be tuned from real behavior. But the sizing model must understand what Kubernetes will actually reserve for scheduling.
Use Kubernetes Requests vs Limits for request and limit tuning.
Calculate the baseline workload before choosing nodes
Do not start with a preferred node plan.
Start with the workload.
Create a table like:
| Workload | CPU request / Pod | Memory request / Pod | Required replicas | Total CPU | Total memory |
|---|---|---|---|---|---|
| Web | measured value | measured value | required floor | calculate | calculate |
| API | measured value | measured value | required floor | calculate | calculate |
| Worker | measured value | measured value | required floor | calculate | calculate |
| System workloads | measured separately | measured separately | per node/cluster | calculate | calculate |
Then add deployment and failure requirements.
This avoids a common mistake: buying a few nodes first and adjusting requests until workloads happen to fit.
Capacity planning should work in the opposite direction.
Size for rollout overlap, not only steady state
A rolling deployment can temporarily run old and new Pods at the same time.
For a Deployment, surge behavior depends on rollout configuration such as maxSurge and maxUnavailable.
That means the cluster may need more schedulable capacity during deployment than during steady state.
If a critical service normally runs four replicas and the rollout policy permits extra replicas, those temporary Pods must still fit somewhere.
Include the actual deployment strategy in sizing.
Ask:
- How many extra Pods can a rollout create?
- Can those Pods fit while all existing replicas are still healthy?
- Does anti-affinity force the new replicas onto different nodes?
- Would the rollout stall because the cluster is too tightly packed?
A cluster that only fits its steady-state workload may be undersized for normal maintenance.
Add explicit failure headroom
Failure reserve should correspond to a named availability target.
For many small production clusters, the most useful test is:
Can the required workload still schedule if one worker is unavailable?
This is often called an N+1 capacity model.
The test is not simply whether total capacity remains above total requests. Placement constraints also matter.
For each worker pool:
- remove the largest or most important eligible worker from the capacity model;
- calculate the remaining allocatable CPU and memory;
- verify that required workload requests can still fit;
- confirm topology, affinity, taints, and storage constraints do not prevent rescheduling.
If required Pods cannot fit, the cluster may be healthy during normal operation but unable to tolerate the node failure or maintenance drain you expect it to survive.
Maintenance drains use the same spare capacity as failures
Planned maintenance looks similar to a temporary node failure from the scheduler's perspective.
When a node is drained, its evictable workloads need another eligible node.
Before calling a cluster production-ready, test:
kubectl drain <node-name> --ignore-daemonsets
Do this only in an appropriate test or maintenance context with a rollback plan.
The operational question is:
Can the cluster drain one worker without leaving required workloads Pending or violating the service's availability objective?
If not, either:
- increase baseline worker capacity;
- reduce or correct workload requests;
- change placement constraints;
- adjust replica design;
- use an autoscaling strategy that can add capacity early enough for the maintenance process.
Do not rely on emergency node creation as the only plan for routine maintenance.
Choose worker size from workload shape
After calculating required capacity, choose worker shapes.
Larger and smaller workers trade different risks.
Larger workers
Advantages:
- fewer nodes to operate;
- less repeated DaemonSet overhead;
- more room for Pods with large requests;
- potentially better bin packing for mixed workloads.
Trade-offs:
- losing one node removes more cluster capacity;
- one large workload can create a larger blast radius;
- scaling occurs in larger capacity increments.
Smaller workers
Advantages:
- smaller failure unit;
- finer-grained scaling;
- workloads can be spread across more nodes.
Trade-offs:
- more per-node overhead;
- more DaemonSet replicas;
- more nodes to manage;
- more potential resource fragmentation.
The useful rule is:
Choose nodes large enough for the largest normal Pod, but not so large that losing one node removes more capacity than your failure model can absorb.
CPU and memory ratios matter more than total capacity alone
A cluster can have enough total resources and still be unable to schedule Pods.
Example:
remaining cluster capacity: 8 vCPU 4 GiB memory new Pod request: 1 vCPU 6 GiB memory
The cluster has plenty of CPU, but the Pod cannot fit because memory is the constrained resource.
The opposite can happen for CPU-heavy workloads.
Track both:
- requested CPU / allocatable CPU;
- requested memory / allocatable memory.
Do not average them into one utilization number.
Also inspect the largest Pod requests. One oversized Pod can determine the minimum node shape even if total cluster utilization is low.
Placement constraints can create local capacity shortages
Total cluster capacity is not always usable by every Pod.
Scheduling can be constrained by:
- nodeSelector;
- node affinity;
- pod anti-affinity;
- taints and tolerations;
- topology spread constraints;
- storage topology;
- architecture requirements;
- dedicated node pools.
A cluster may have 20 GiB of free memory overall while a specific workload has no eligible node with 4 GiB free.
For this reason, capacity planning should be performed per eligible worker pool, not only across the whole cluster.
Use Kubernetes Node Pools when workloads require different scaling, isolation, or capacity classes.
Pod density is a ceiling, not a target
Kubernetes v1.37 documents large-cluster configurations of up to:
- 110 Pods per node;
- 5,000 nodes;
- 150,000 total Pods;
- 300,000 total containers.
These are scalability limits under validated large-cluster conditions, not recommended small-team utilization targets.
A production node can reach CPU, memory, network, storage, or operational limits long before it reaches the maximum Pod count.
Pod density is also affected by:
- per-Pod networking overhead;
- logging;
- image pulls;
- DaemonSets;
- container count;
- ephemeral storage;
- application startup behavior.
Size from workload resources and operational behavior first.
Headroom should have a reason
There is no universal Kubernetes headroom percentage that fits every cluster.
Instead, reserve capacity for specific requirements:
| Headroom | What it protects |
|---|---|
| Node-failure reserve | Required workloads still fit after one worker is lost |
| Maintenance reserve | Workloads can move during drains/upgrades |
| Rollout reserve | Surge replicas fit during deployments |
| Autoscaler warm-up | Existing workers absorb short-term demand while new nodes join |
| Growth reserve | Known near-term workload growth does not force emergency changes |
Avoid an unexplained rule such as "always leave 30% free."
A lower reserve may be acceptable for a development environment.
A customer-facing workload with strict availability may need more.
Document the reason for the reserve and review it as workload behavior changes.
Autoscaling does not replace the safe minimum fleet
Autoscaling is a capacity mechanism, not instant failover.
When Pods become unschedulable, adding a node still takes time:
- the scheduler identifies Pending Pods;
- the autoscaler decides additional capacity is needed;
- infrastructure is provisioned;
- the node joins the cluster;
- system workloads start;
- the node becomes Ready;
- workloads are scheduled.
Your minimum worker pool should therefore be able to support the workload conditions that must survive immediately.
Use autoscaling for additional demand above the safe baseline.
Define:
- minimum node count;
- maximum node count;
- expected worker shape;
- scale-up triggers;
- scale-down policy;
- acceptable provisioning latency;
- maximum acceptable monthly capacity.
Use Kubernetes Autoscaling for the deeper interaction between HPA and node scaling.
Minimum cluster size depends on the availability target
There is no universal number of worker nodes for every Kubernetes cluster.
A development cluster may intentionally accept a small topology.
A production cluster should derive node count from:
- required replicas;
- failure tolerance;
- placement rules;
- maintenance model;
- worker shape;
- expected traffic;
- autoscaling behavior.
For Raff Managed Kubernetes specifically, the current product maintains a two-worker floor for autoscaled node pools.
That platform constraint is not a universal Kubernetes recommendation. It is a Raff product behavior and should be combined with the workload's own availability target.
If your application must remain healthy after one worker is unavailable, verify the remaining worker capacity—not simply the node count.
Control-plane sizing is separate from worker sizing
In self-managed Kubernetes, the team must size and operate the control plane, including the API server and datastore.
In managed Kubernetes, worker sizing can be handled independently.
Raff currently offers:
- a $0 standard managed control plane;
- optional three-master HA for $30/month.
Worker resources are billed separately and should be sized from workload demand.
For production workloads, control-plane HA is an availability decision. It does not add application worker capacity.
Use Managed vs Self-Managed Kubernetes for the ownership trade-off.
Current Raff worker shapes
Current Raff Kubernetes worker plans include:
| Worker | vCPU | Memory | Monthly price |
|---|---|---|---|
| Starter | 1 | 2 GB | $9.99 |
| Standard | 2 | 4 GB | $17.99 |
| Performance | 4 | 8 GB | $33.99 |
| High Memory | 8 | 16 GB | $71.99 |
| Large | 8 | 32 GB | $119.99 |
| Scale | 16 | 64 GB | $229.99 |
Dedicated Kubernetes storage is priced separately at $0.10/GB-month, from 10 GB to 1,000 GB per storage node.
Use the Raff Kubernetes product page for live pricing and product constraints because worker plans can change over time.
The correct workflow is still:
calculate required capacity first → choose worker shapes second → calculate cost third.
Do not reverse that order.
A practical sizing worksheet
Use this sequence for a new or existing cluster.
Step 1 — Inventory minimum workloads
List every production workload and its minimum replica count.
Step 2 — Sum CPU and memory requests
Calculate requested CPU and memory for the minimum healthy workload set.
Step 3 — Measure platform overhead
Use real Node Allocatable and system Pod requests where possible.
Step 4 — Add rollout overlap
Account for surge replicas and maintenance behavior.
Step 5 — Define failure tolerance
Decide whether one worker or another named failure must be tolerated without waiting for new capacity.
Step 6 — Apply placement constraints
Split the model by node pool or eligibility class.
Step 7 — Choose worker shapes
Select nodes that fit both aggregate capacity and the largest individual Pods.
Step 8 — Verify N+1 capacity
Remove one worker from the model and recalculate.
Step 9 — Set autoscaling boundaries
Choose safe minimums and explicit maximums.
Step 10 — Validate with production data
After launch, compare:
- requests vs actual usage;
- Pending Pod events;
- deployment surge behavior;
- node drains;
- autoscaler activity;
- worker utilization;
- cost.
Sizing is a recurring capacity decision, not a one-time launch calculation.
Common Kubernetes sizing mistakes
Using utilization instead of requests
Observed utilization alone does not tell you whether the scheduler can place another Pod.
Using requests without validating them
Bad requests produce bad sizing. Review requests against actual workload behavior.
Ignoring Node Allocatable
Raw node capacity overstates what application Pods can use.
Ignoring deployment surge
Clusters that fit steady state can still fail during rolling updates.
Ignoring failure domains
Total capacity can look sufficient while one worker pool lacks the resources to tolerate a node loss.
Letting autoscaling define the minimum
Autoscaling is not instant. Establish a safe baseline first.
Buying node shapes before calculating workloads
Capacity requirements should determine node selection, not the other way around.
Final sizing rule
Start with workload requests and measured Node Allocatable.
Add system overhead, rollout overlap, a named failure reserve, and justified growth headroom. Then choose worker shapes that fit both the aggregate demand and the largest individual Pods. Finally, test whether the cluster still meets its availability objective when one eligible worker is unavailable.
That is a defensible Kubernetes sizing model.
Continue with Kubernetes Node Pools for workload separation, Kubernetes Requests vs Limits for request tuning, Kubernetes Autoscaling for dynamic capacity, and Kubernetes Cluster Management for the full day-2 operating model.