Kubernetes upgrade strategy is the repeatable process for moving a cluster to a supported version while preserving workload availability, API compatibility, recovery options, and operational control.
For small teams, the safest model is:
patch routinely → plan minor upgrades → validate compatibility → protect recovery state → upgrade control plane → rotate workers → verify workloads → close only after user-facing checks pass
Raff Technologies now supports in-place Kubernetes upgrades from the dashboard and API, rolling one node at a time while the cluster API remains reachable. Raff also supports automatic patch upgrades inside a maintenance window or fully manual upgrades.
Managed upgrades reduce the platform work, but the application team still owns workload compatibility, PodDisruptionBudgets, recovery, stateful behavior, and acceptance testing.
This guide owns the Kubernetes upgrade decision model. It does not replace provider-specific click paths or self-managed kubeadm procedures.
Keep the cluster on a supported Kubernetes branch
Kubernetes currently maintains the most recent three minor release branches.
As of September 2026, those branches are:
- Kubernetes 1.37
- Kubernetes 1.36
- Kubernetes 1.35
Kubernetes 1.19 and newer generally receive about one year of patch support.
This creates a real operational deadline.
A cluster can continue running after a minor branch reaches end of life, but the upstream project no longer provides normal bug and security fixes for that branch.
Version age is therefore operational debt.
Do not wait until the current branch is nearly unsupported before discovering:
- removed APIs;
- incompatible operators;
- old admission webhooks;
- unsupported CSI/CNI components;
- outdated Helm charts;
- kubectl/tooling skew;
- workload disruption problems.
A small team should always know:
- current Kubernetes version;
- latest supported patch in that minor;
- next target minor;
- known blockers;
- owner of the upgrade;
- planned validation environment;
- recovery state.
Treat patch and minor upgrades differently
Not every Kubernetes change needs the same process.
| Upgrade type | Default operating model |
|---|---|
| Patch update | Routine maintenance |
| Minor upgrade | Planned compatibility and maintenance campaign |
| Critical security fix | Accelerated review and maintenance |
| End-of-support pressure | High-priority lifecycle work |
Kubernetes recommends applying current patch releases promptly.
For minor upgrades, do more than change the version number.
Review:
- API compatibility;
- workloads;
- controllers;
- CRDs;
- webhooks;
- CNI;
- CSI;
- Gateway/Ingress controllers;
- monitoring;
- backup tooling;
- kubectl and automation.
The maintenance window should be the execution phase, not the discovery phase.
Upgrade one minor version at a time
Kubernetes component upgrade rules do not support skipping minor versions for the normal control-plane upgrade path.
A safe pattern is:
1.35.latest ↓ 1.36.latest ↓ validate ↓ 1.37.latest
If the cluster is several versions behind, plan sequential upgrades.
Before moving to the next minor, Kubernetes recommends being on the latest patch of the current minor and then upgrading to the latest patch of the target minor.
This reduces the chance that the team carries already-fixed defects into the next version transition.
Understand version skew before the change window
During an upgrade, not every component changes at exactly the same moment.
Kubernetes defines supported skew.
Current upstream rules include:
- HA kube-apiserver instances may differ by at most one minor version;
- kubelet must not be newer than kube-apiserver;
- kubelet may be up to three minor versions older than kube-apiserver;
- kube-proxy may be up to three minor versions older than kube-apiserver;
- kubectl is supported within one minor version older or newer than kube-apiserver.
These ranges are transition allowances.
Do not use them as a reason to leave a cluster permanently split across versions.
The safest steady state is still a cluster whose components are intentionally aligned.
Control plane first, workers after
At a high level, Kubernetes documents this upgrade order:
- control plane;
- worker nodes;
- clients such as kubectl;
- manifests/resources that need adjustment for API changes.
In managed Kubernetes, the provider may execute much of this sequence.
The workload team still needs to understand the order because it determines:
- version-skew expectations;
- when workers are drained;
- when application disruption can occur;
- when acceptance testing should happen.
Do not treat "managed" as "no maintenance model required."
Review APIs and controllers before production
A minor upgrade can break the cluster even when application images do not change.
Before upgrading, review:
- Kubernetes release notes;
- deprecated and removed APIs;
- CustomResourceDefinitions;
- operators/controllers;
- admission webhooks;
- CNI and NetworkPolicy implementation;
- CSI/storage drivers;
- Gateway API or Ingress controllers;
- metrics and logging components;
- backup/restore tooling;
- Helm charts;
- CI/CD;
- Terraform/provider versions;
- kubectl.
Pay special attention to resources that are installed once and then rarely revisited.
Application Deployments are often actively maintained.
Old webhooks, operators, and platform add-ons are more likely to become hidden upgrade blockers.
Test the target version before production
Production should not be the first environment to discover an upgrade incompatibility.
A representative validation environment should prove that the target version can:
- start critical controllers;
- schedule workloads;
- mount storage;
- resolve DNS;
- route traffic;
- pass readiness probes;
- run Jobs and CronJobs;
- support current CI/CD;
- restore required data;
- expose expected metrics/logs.
The test environment does not need production traffic volume.
It needs the components and behaviors that could fail because the Kubernetes version changed.
Worker maintenance requires schedulable headroom
Worker upgrades require workload movement.
Kubernetes recommends draining a node before a minor-version kubelet upgrade.
That means a production cluster needs somewhere for the displaced workload to go.
Before the upgrade, test:
Can one worker be unavailable while required workloads remain healthy?
Use Kubernetes Cluster Sizing if this is not clear.
Check:
- allocatable CPU and memory on remaining workers;
- resource requests;
- anti-affinity;
- topology spread;
- node selectors;
- taints/tolerations;
- storage topology;
- rollout surge;
- startup time.
An upgrade is not safe because the total cluster has enough CPU.
It is safe only if the remaining eligible nodes can schedule the workloads that must survive.
PodDisruptionBudgets protect voluntary disruption, not capacity
PodDisruptionBudgets help during voluntary maintenance such as node drains.
They can limit how many replicas are disrupted at once.
They do not:
- create spare worker capacity;
- protect against involuntary node failure;
- fix an application with one replica;
- override impossible scheduling constraints.
Example:
3 replicas PDB requires 2 available 1 worker drained replacement Pod must schedule elsewhere
If the remaining cluster has no eligible capacity, the drain can stall.
Therefore PDB design and capacity planning must be reviewed together.
Readiness and graceful shutdown matter during drains
A rolling worker upgrade depends on application lifecycle behavior.
Before maintenance, verify:
- readiness removes unhealthy/new Pods from traffic correctly;
- terminationGracePeriodSeconds is appropriate;
- the application handles SIGTERM;
- in-flight work can finish or retry safely;
- queue workers do not lose acknowledged jobs;
- stateful workloads can detach and reattach storage;
- connections can move without corrupting state.
A Pod moving successfully is not enough.
The workload must remain functionally correct during movement.
Raff now supports in-place rolling Kubernetes upgrades
Raff's current managed Kubernetes platform supports version upgrades in place from the dashboard or API.
The platform performs rolling upgrades one node at a time while keeping the cluster API reachable.
Raff also supports:
- automatic patch upgrades in a maintenance window;
- fully manual upgrade control;
- control-plane health visibility;
- node removal that drains workloads first;
- PodDisruptionBudget-aware node removal.
These features reduce the amount of lifecycle machinery a small team must operate itself.
They do not remove workload-level readiness requirements.
Before starting an upgrade on Raff, the application team should still verify:
- workload API compatibility;
- required replicas;
- PDB behavior;
- spare worker capacity;
- stateful workload recovery;
- add-on compatibility;
- acceptance tests;
- business-data backups.
Maintenance windows need stop criteria
An upgrade window should define not only when the change starts but also when the team stops.
Write down abort criteria before maintenance.
Examples:
- API error rate exceeds threshold;
- critical workload cannot become Ready;
- stateful workload cannot reattach storage;
- DNS or service discovery fails;
- public traffic path is broken;
- control-plane health degrades;
- scheduler cannot place required replicas;
- backup/recovery state is no longer trustworthy.
Do not decide the acceptable failure threshold after the failure occurs.
A maintenance window is safer when the team knows:
- what is expected;
- what is abnormal;
- who makes the stop decision;
- what recovery path follows.
Separate application rollback from Kubernetes recovery
"Rollback" can mean several different things.
Application rollback
Revert:
- container image;
- Deployment configuration;
- Helm release;
- application feature flag.
This may be possible even when the cluster remains on the new Kubernetes version.
Add-on rollback
A controller, chart, CNI, CSI, or Gateway/Ingress component may need its own version rollback.
Again, that does not necessarily imply reversing the Kubernetes cluster version.
Data recovery
A database or PersistentVolume may require restoration from a protected recovery point.
That is a data-protection operation.
Cluster-version recovery
Returning the entire cluster to an older Kubernetes version is provider/distribution-specific and should not be treated as an assumed in-place downgrade.
Do not build the maintenance plan around:
"If it fails, we'll just downgrade Kubernetes."
Instead, define what can be repaired forward, what can be rolled back independently, and what provider-specific cluster recovery capability exists.
Recovery state must exist before the upgrade
Before a discretionary minor upgrade, know:
- latest usable application backup;
- database recovery point;
- persistent-storage recovery method;
- manifests / Git revision for the known-good application state;
- previous compatible controller/add-on versions;
- current cluster configuration;
- owner of recovery;
- expected restore time.
Use Kubernetes Backup and Disaster Recovery Strategy for the full recovery model.
Replicated storage and PDBs are not backups.
A healthy HA cluster can still contain corrupted business data.
Validate each upgrade layer separately
After the control plane changes, verify the control plane.
After workers change, verify worker and scheduling behavior.
After the whole cluster is upgraded, verify applications.
Control plane
Check:
- API availability;
- controller health;
- scheduler health;
- etcd/control-plane health where exposed;
- recurring API-server errors.
Raff now exposes control-plane health signals including etcd latency, component restarts, and API-server errors.
Worker layer
Check:
kubectl get nodes kubectl get pods -A
Confirm:
- nodes are Ready;
- expected version is present;
- workloads can schedule;
- no unexpected Pending Pods remain;
- network and storage components are healthy.
Application layer
Verify:
- critical user paths;
- error rate;
- latency;
- Jobs/workers;
- authentication;
- database reads/writes;
- external integrations.
The upgrade is not complete when all nodes report the new version.
It is complete when the platform and the application are both operating normally.
Use an explicit Kubernetes upgrade checklist
Before the window
- Current cluster version recorded.
- Target version is supported.
- Current minor is on an appropriate current patch.
- No unsupported minor-version skip is planned.
- Release notes reviewed.
- Deprecated/removed APIs checked.
- CRDs/operators/webhooks reviewed.
- CNI/CSI/Gateway/Ingress components reviewed.
- kubectl/CI/CD/Terraform compatibility checked.
- Workloads tested against target version.
- Recovery state verified.
- One-worker maintenance headroom tested.
- PDB behavior reviewed.
- Stop criteria and owner documented.
During the window
- Control plane health monitored.
- Worker changes proceed gradually.
- Drain behavior is observed.
- Required workloads remain available.
- Unexpected Pending Pods are investigated before continuing.
- Stateful workloads are checked after movement.
- Upgrade is paused on unexplained platform degradation.
After the window
- Expected Kubernetes version confirmed.
- All required nodes are Ready.
- Critical controllers are healthy.
- DNS/networking works.
- Persistent storage works.
- Public traffic works.
- Application acceptance tests pass.
- Error and latency signals are normal.
- Backup/recovery assumptions remain valid.
- Upgrade result is recorded.
A small-team upgrade cadence
A practical policy is:
Continuously
Watch Kubernetes release and security information.
Patch releases
Review and apply through routine maintenance, especially for important security/reliability fixes.
Minor releases
Plan a compatibility campaign before the current branch approaches end of support.
Before every upgrade
Validate compatibility, disruption capacity, and recovery.
After every upgrade
Record:
- version;
- issues found;
- duration;
- workload impact;
- recovery actions;
- follow-up fixes.
This turns maintenance into a repeatable system instead of institutional memory.
How managed Kubernetes changes the upgrade workload
Self-managed Kubernetes requires the team to operate:
- control-plane upgrades;
- datastore lifecycle;
- certificates;
- component sequencing;
- node lifecycle;
- recovery mechanics.
Managed Kubernetes moves much of that platform machinery to the provider.
On Raff, in-place version upgrades, maintenance-window patching, rolling worker changes, PDB-aware drains, and control-plane health reduce that operational burden.
The application team still owns the parts closest to customers:
- API compatibility;
- workload disruption;
- persistent data;
- acceptance tests;
- observability;
- application rollback;
- business recovery.
That is the useful boundary.
Managed Kubernetes does not make upgrades irrelevant.
It makes the team's upgrade work more focused.
Final recommendation
Keep Kubernetes upgrades routine.
Stay on supported branches, apply patch updates without unnecessary delay, and treat minor upgrades as planned compatibility changes.
Do not begin a production upgrade until:
- the target version is supported;
- APIs and add-ons are compatible;
- workloads can survive worker maintenance;
- required data is recoverable;
- stop criteria are written;
- application acceptance tests exist.
Then upgrade in controlled stages and confirm recovery with the same user-facing signals that define normal service health.
Continue with Kubernetes Cluster Management, Kubernetes Cluster Sizing, Kubernetes Monitoring, and Kubernetes Backup and Disaster Recovery for the operating systems around maintenance.