A K3s upgrade changes more than one binary. It changes the Kubernetes version and the K3s-packaged control plane, runtime, networking, ingress, storage helpers, and other bundled components that your cluster depends on.
For a self-managed production cluster, the safest K3s upgrade has four gates:
- choose a supported target version deliberately;
- prove that cluster state can be restored before changing the first server;
- upgrade server nodes first, one at a time, then agent nodes; and
- define rollback as a datastore recovery operation, not simply “install the old binary again.”
September 2026 update: K3s v1.37 moves the embedded HA datastore to etcd 3.7.1 and updates several packaged components, including Traefik and Helm-controller behavior. Treat a cross-minor upgrade as a cluster-state change with a preserved pre-upgrade recovery point, not as an ordinary package refresh.
K3s supports both manual upgrades and Kubernetes-native automated upgrades through Rancher's System Upgrade Controller. Both approaches can be safe. The better choice depends on cluster size, how much rollout automation the team wants, and whether the maintenance process is already declarative.
This guide owns the K3s-specific upgrade and rollback model: release channels, version pinning, manual upgrades, System Upgrade Controller Plans, server/agent sequencing, datastore recovery gates, and K3s rollback. For generic Kubernetes API deprecation, version-skew, PodDisruptionBudget, and platform-wide compatibility planning, use Kubernetes Upgrade Strategy: Maintenance, Risk, and Rollback.
K3s upgrade strategy in one table
| Decision | Recommended K3s baseline | Why |
|---|---|---|
| Production release source | stable channel or an explicitly pinned tested version | K3s recommends stable for production |
| Minor-version path | Do not skip intermediate Kubernetes minor versions | Kubernetes version-skew policy still applies |
| Upgrade order | Servers first, one at a time; agents afterwards | Protects control-plane compatibility and HA |
| Embedded-etcd servers | Keep quorum throughout maintenance | Losing quorum turns an upgrade into a datastore incident |
| Recovery point | Create/verify datastore backup before first server change | Previous-minor rollback requires older-minor datastore state |
| Server token | Back it up with the datastore | It is required to decrypt bootstrap data during restore |
| Manual upgrade | Re-run install script with the original configuration or replace binary | Simple for small clusters |
| Automated upgrade | System Upgrade Controller with separate server and agent Plans | Better repeatability for multi-node clusters |
| Rollback | Stop rollout, restore old datastore state when required, then run matching K3s version | Downgrade is not a normal rolling upgrade |
The operating principle is straightforward:
An upgrade is reversible only if the team has preserved the state needed by the version it intends to return to.
Manual and automated K3s upgrades solve different operating needs
K3s officially supports two main upgrade approaches.
Manual upgrade
The operator upgrades each node directly, either by:
- re-running the K3s installation script; or
- replacing the K3s binary and restarting the service.
This is often the cleanest approach for:
- a single-server cluster;
- a three-server HA cluster maintained infrequently;
- teams that want explicit control over every node change;
- infrastructure managed through an external automation system.
System Upgrade Controller
The System Upgrade Controller runs inside Kubernetes and uses Plan custom resources to decide:
- which nodes are eligible;
- which K3s version or release channel they should target;
- how many nodes can be changed at once;
- whether nodes are cordoned;
- when server upgrades must finish before agent upgrades start;
- which maintenance window may create upgrade Jobs.
This is useful when:
- several nodes need the same lifecycle policy;
- upgrades should be declarative and repeatable;
- server and worker sequencing should be encoded rather than remembered;
- maintenance windows need an explicit schedule.
| Factor | Manual K3s upgrade | System Upgrade Controller |
|---|---|---|
| Best cluster size | Small | Small to larger multi-node fleets |
| Node sequencing | Operator-managed | Plan-managed |
| Version target | Script/binary version or channel | version or channel in Plan |
| Server-before-agent dependency | Manual procedure | Can be encoded with separate Plans |
| Concurrency | Operator-controlled | concurrency field |
| Maintenance window | External process | Plan window can restrict Job creation |
| Rollout visibility | Service/node checks | Plans + Jobs + nodes |
| Main risk | Human inconsistency | A bad Plan can automate the wrong change |
Automation should reduce repeated work, not remove the approval gate before production maintenance.
Choose the target release before opening the maintenance window
K3s supports release channels for both installation-script and automated upgrades.
Current official channels include:
stable— the default and K3s-recommended production channel; releases have had a period of community testing;latest— the highest non-prerelease version by semantic version ordering, without the same community-testing period;- minor-version channels, such as
v1.35— resolve to the newest K3s release for that Kubernetes minor, which is not necessarily a release promoted tostable.
This creates three useful production policies.
Follow stable automatically
A team can point a System Upgrade Controller Plan at the stable channel and allow the cluster to follow new stable K3s releases.
That reduces version-maintenance effort, but it also means a new target can become eligible when the channel moves.
Use this only when:
- maintenance windows are controlled;
- backups are continuous and verified;
- release notes are reviewed as part of the operating process;
- the team accepts automatic eligibility for the next stable release.
Pin a specific K3s version
A fixed version gives the most deterministic change record.
For example, the installation script supports:
INSTALL_K3S_VERSION=vX.Y.Z+k3s1
System Upgrade Controller Plans can also specify an exact version instead of a channel.
This is a strong fit when production promotion follows:
Test cluster ↓ Staging cluster ↓ Pinned production version
Follow a minor channel
A minor channel can keep a cluster on one Kubernetes minor while receiving newer K3s builds from that line.
That can be useful when the team wants patch movement without crossing the next minor boundary automatically.
However, K3s explicitly notes that a minor channel selects the latest release for that minor and that release is not necessarily stable. Treat the channel as a version-selection mechanism, not as a quality guarantee.
Do not skip Kubernetes minor versions
K3s packages Kubernetes, so Kubernetes version-skew rules still apply.
The current K3s upgrade documentation explicitly warns operators not to skip intermediate Kubernetes minor versions. The System Upgrade Controller does not protect the cluster from an unsupported version jump.
If a cluster is several minors behind, plan sequential changes:
v1.N.x ↓ v1.N+1.x ↓ validate v1.N+2.x ↓ validate current target
Do not compress several years of platform lifecycle work into one large maintenance command.
The broad compatibility review—deprecated APIs, CRDs, controllers, CNI, CSI, admission webhooks, and workload behavior—belongs in the Kubernetes Upgrade Strategy. The K3s-specific rule here is simply that the distribution does not make unsupported version jumps safe.
K3s v1.37 adds an important datastore and packaged-component checkpoint
K3s v1.37 was released on September 10, 2026. For upgrade planning, two changes deserve explicit attention.
First, embedded HA clusters move to the etcd 3.7 series, with K3s v1.37 targeting etcd 3.7.1. That makes the datastore version part of the maintenance boundary. Before crossing into v1.37 on embedded etcd, keep a verified snapshot from the older K3s/Kubernetes minor and preserve the matching server token until the rollback window has closed.
Second, K3s v1.37 changes packaged component behavior. The release moves the packaged Helm stack to Helm V4-era behavior and changes the Traefik chart failure policy toward retrying reconciliation instead of reinstalling resources on minor upgrade conflicts. That is intended to make packaged-component upgrades safer, but it does not remove the need to validate ingress, Helm-managed add-ons, and workload traffic after the node upgrade.
For a v1.37 maintenance window, add these checks:
- verify the current datastore type and exact K3s version before changing the first server;
- create and identify the pre-upgrade datastore recovery point;
- retain the matching server token;
- review the K3s v1.37 release notes for packaged-component changes;
- validate embedded-etcd health after every upgraded server;
- validate Traefik/ingress and other packaged Helm components after the server set is complete;
- keep the older-minor snapshot through the agreed rollback window.
This does not mean every v1.37 upgrade should be rolled back at the first error. It means the team should preserve the only state that makes a previous-minor rollback possible before the upgrade starts.
The recovery gate comes before the first server upgrade
The most important pre-upgrade task is not downloading K3s. It is proving that the current cluster can be recovered.
The required backup depends on the K3s datastore.
| K3s datastore | Pre-upgrade recovery material |
|---|---|
| SQLite | Copy of the K3s SQLite datastore + server token |
| Embedded etcd | Valid etcd snapshot + server token |
| External PostgreSQL/MySQL/MariaDB/etcd | Database-native backup/snapshot + server token |
K3s documentation requires the server token at:
/var/lib/rancher/k3s/server/token
The token is not merely a join password. K3s uses it to protect confidential bootstrap data inside the datastore. A datastore backup without the matching token can be unusable during restore.
For embedded etcd, K3s supports an on-demand snapshot:
k3s etcd-snapshot save
Current K3s defaults also schedule embedded-etcd snapshots every 12 hours and retain 5 scheduled snapshots, but a production upgrade should not assume that the default snapshot happens to be the recovery point the team needs.
Record the exact snapshot that represents the pre-upgrade state and verify that it can be accessed from the recovery environment.
Cluster-state backup is not application-data backup
A K3s datastore backup protects Kubernetes cluster state. It does not automatically protect the bytes stored behind every PersistentVolumeClaim.
Before maintenance, separate these layers:
K3s datastore recovery → API objects, cluster state, Secrets, metadata PVC / storage recovery → application files and volume data Database recovery → transactionally consistent business data
If the upgrade affects storage components or stateful applications, verify the application-data recovery path too.
Use K3s Persistent Storage and Backup Strategy for the storage boundary and K3s High Availability: Embedded etcd vs External Database for datastore-specific HA and backup ownership.
At Raff, we treat this as two upgrade gates: first prove the control plane and datastore can recover; then prove workload nodes can rotate without creating a second incident. If the team cannot name the usable snapshot and matching server token before the first server changes, the maintenance window is not ready.
Manual K3s upgrades require the original configuration
K3s allows an existing node to be upgraded by re-running the installation script.
A production operator should know one important behavior: the installer regenerates the systemd or OpenRC service configuration from the values supplied to it.
Current K3s documentation warns that if the original installation used:
INSTALL_K3S_EXEC;K3S_...environment variables; or- trailing shell arguments,
those values must be supplied again when the installer is re-run. If they are omitted, the original values can be lost.
A safer long-term pattern is to keep persistent K3s configuration in the K3s configuration file, because the install script does not manage the contents of that file.
Before a manual upgrade, therefore document:
current K3s version + current config file + installer environment + installer arguments + systemd/openrc service state
Then choose either a channel or exact version and run the same node configuration against the intended K3s release.
This detail is easy to miss and can turn a version upgrade into an accidental configuration change.
Server nodes always come before agent nodes
K3s's current documentation is explicit:
upgrade server nodes first, one at a time, then agent nodes.
That order gives the cluster a predictable control-plane compatibility path.
A simplified sequence is:
Server 1 ↓ validate Server 2 ↓ validate Server 3 ↓ validate control plane Agent 1 ↓ Agent 2 ↓ Agent N ↓ validate workloads
Do not start upgrading workers because they appear less critical. Agents depend on the control plane, and the intended K3s sequence upgrades the server side first.
Single-server K3s has the simplest upgrade path and the largest control-plane interruption
A single-server cluster has no second API server to carry the control plane during a restart.
The upgrade pattern is therefore simple:
1. Verify current version and recovery material 2. Protect application state where required 3. Upgrade the single K3s server 4. Verify K3s service and API 5. Verify system Pods, ingress, storage, and applications
K3s notes that application containers continue running when the K3s service is stopped, and a normal K3s restart does not automatically drain the node.
For workloads that are sensitive to a brief API-server outage or node-level disruption, explicitly cordon/drain according to the workload's maintenance design before changing the server.
A single-server upgrade can be operationally easier than HA maintenance, but it also has no redundant control plane if the upgraded server fails to return.
That makes the pre-upgrade recovery gate especially important.
Embedded-etcd HA upgrades are quorum maintenance
In a three-server embedded-etcd K3s cluster, each server is both a control-plane node and an etcd member.
A safe upgrade sequence keeps quorum intact:
Server A upgrade ↓ healthy + Ready Server B upgrade ↓ healthy + Ready Server C upgrade ↓ healthy + Ready Agents afterwards
With three etcd members, the cluster needs 2 available members for quorum.
Do not:
- upgrade multiple server nodes simultaneously;
- reboot two servers together;
- combine an upgrade with unrelated etcd/storage maintenance;
- continue to the next server while the previous member is unhealthy;
- assume that healthy application Pods prove etcd is healthy.
For a small HA cluster, concurrency: 1 is the conservative System Upgrade Controller setting for the server Plan and matches the official K3s example.
The correct stop condition is not “the Job completed.” It is “the upgraded server has rejoined the healthy control plane and datastore before the next member changes.”
External-datastore K3s has a different recovery gate
When K3s uses PostgreSQL, MySQL, MariaDB, or external etcd, the datastore lifecycle is outside K3s.
The K3s server sequence is still:
server nodes first one at a time agents afterwards
But the pre-upgrade backup must come from the external datastore's own recovery system.
That means the maintenance owner should know:
- which database snapshot/dump corresponds to the current K3s version;
- how long restoration takes;
- whether the datastore endpoint can be restored without changing K3s configuration;
- whether credentials and TLS materials remain valid;
- who owns the database recovery decision.
The detailed embedded-etcd versus external-database choice belongs in the K3s HA guide. Upgrade strategy only needs to respect the topology already chosen.
System Upgrade Controller makes the sequence declarative
The K3s System Upgrade Controller watches Plan resources and creates upgrade Jobs on matching nodes.
K3s recommends at least two Plans:
- a server Plan targeting control-plane/server nodes; and
- an agent Plan targeting nodes without the control-plane label.
The current official example uses:
concurrency: 1for the server Plan;cordon: true;- a control-plane node selector for the server Plan;
- an agent selector that excludes control-plane nodes;
- a
preparestep on the agent Plan that waits for the server Plan to finish.
Conceptually:
server-plan ├── server 1 ├── server 2 └── server 3 ↓ complete agent-plan ├── agent 1 ├── agent 2 └── agent N
The agent dependency is implemented by the K3s upgrade image's prepare logic; it is not a generic System Upgrade Controller guarantee for arbitrary Plans.
That distinction matters when teams customize the example.
Pin a Plan version when production approval must be explicit
A Plan can target either:
- a
channel; or - a fixed
version.
A channel is convenient for continuous maintenance. A version is easier to audit as a single approved production change.
For teams with a staging gate, a fixed-version pattern is often easier to reason about:
Staging validates vX.Y.Z+k3s1 ↓ Production Plans are changed to vX.Y.Z+k3s1 ↓ Server Plan completes ↓ Agent Plan completes
Do not mix latest with the assumption that “latest” means “production recommended.” K3s explicitly distinguishes latest from stable.
Maintenance windows can be encoded in Plans
Current System Upgrade Controller Plans support a window field.
The window can restrict when new upgrade Jobs are created by:
- day;
- start time;
- end time; and
- timezone.
That is useful for production because an always-active Plan does not have to mean an always-open maintenance window.
One important limitation: a Job that starts during the allowed window may continue running after the window closes.
So the window is a Job creation boundary, not a guaranteed completion deadline.
Operationally, the team still needs enough time after the final expected node change for validation and recovery.
Upgrade Jobs are highly privileged
System Upgrade Controller needs deep node access to replace the K3s binary and restart underlying services.
K3s documentation notes that the upgrade Job is highly privileged and, by default, uses host namespaces, CAP_SYS_BOOT, and a read/write mount of the host root.
Treat the controller and its Plans as infrastructure-administration capability.
That means:
- restrict who can create or modify
Planresources; - review the controller manifest and image source;
- do not allow application namespaces to manage upgrade Plans;
- monitor changes to
system-upgraderesources; - remove abandoned Plans that no longer represent the intended policy.
An automated upgrade controller is part of the cluster's privileged control surface.
Monitor the Plan, Jobs, nodes, and workloads together
K3s documents the following commands for upgrade progress:
kubectl -n system-upgrade get plans -o wide kubectl -n system-upgrade get jobs
Those show the orchestration layer, but production validation should also include cluster and workload state.
After each server, and again after the full server set, verify:
kubectl get nodes -o wide kubectl get pods -A
Also confirm:
- the expected K3s/Kubernetes version is reported;
- server nodes are
Ready; - embedded-etcd quorum remains healthy when used;
- CoreDNS and cluster networking work;
- Traefik or the chosen ingress layer is healthy;
- ServiceLB or external exposure still behaves correctly;
- storage controllers and mounted PVCs work;
- critical controllers are not crash-looping;
- application health checks and real user paths pass.
Use K3s Networking, Ingress, and Load Balancing when the post-upgrade failure is actually an ingress, CNI, ServiceLB, or API-endpoint problem rather than the K3s binary itself.
Stop the rollout before deciding to roll back the cluster
Not every failed node upgrade requires a full cluster rollback.
Use a failure hierarchy.
| Failure | First response | Full K3s rollback? |
|---|---|---|
| One agent upgrade Job fails | Stop/repair agent rollout | Usually no |
| One server fails but HA quorum remains | Stop before changing another server; recover that member | Not automatically |
| CNI/ingress add-on incompatibility | Stop rollout; assess component/version fix | Maybe |
| Application incompatible but cluster healthy | Roll back application/config | Usually no |
| API/control plane unusable after version change | Enter control-plane recovery procedure | Possibly |
| Datastore state incompatible with previous minor | Restore older-minor datastore snapshot as part of rollback | Yes for previous-minor rollback |
A rollback is a higher-risk recovery action than stopping a rollout.
The safest maintenance process therefore has separate decisions for:
- pause — do not upgrade another node;
- repair forward — fix the changed node while staying on the new version;
- application rollback — revert workload changes without changing K3s;
- cluster rollback — return K3s and datastore state to an older version.
System Upgrade Controller intentionally prevents downgrades
This is one of the strongest reasons to define rollback before using automated upgrades.
K3s documentation states that Kubernetes does not support downgrading control-plane components. The k3s-upgrade image used by System Upgrade Controller therefore refuses to downgrade K3s.
If a Plan points to a version lower than a node's current K3s version, the Plan fails.
If cordon: true was configured, affected nodes can remain cordoned after that failure.
The documented recovery choices are to:
- change the Plan back to the same or a newer valid version so it can succeed; or
- delete the failed Plan and manually uncordon the affected nodes.
Do not attempt to express a production rollback by simply editing the Plan to an older version.
Real K3s rollback combines binary downgrade and datastore restoration
K3s now documents an explicit rollback procedure.
A previous-minor rollback is not a rolling downgrade.
K3s states that rollback uses a combination of:
- running the older K3s binary; and
- restoring datastore state compatible with the version being restored.
Most importantly:
When rolling back to a previous Kubernetes minor version, K3s requires a datastore snapshot taken while the cluster was running that older minor. If the database cannot be restored, the cluster cannot roll back to that previous minor.
This is why the pre-upgrade recovery point is part of the change plan rather than a generic backup checkbox.
The rollback procedure differs by datastore:
| Datastore | Rollback foundation |
|---|---|
| SQLite | Restore the backed-up SQLite database and matching K3s version |
| Embedded etcd | Stop cluster processes, use older K3s binary, restore older snapshot, rebuild etcd membership, restart servers/agents |
| External database | Stop cluster processes, restore database snapshot, run matching older K3s binary, restart nodes |
Because a rollback intentionally restores historical cluster state, it can also discard Kubernetes changes made after the snapshot.
Application databases and external business data may have continued changing during that time. The recovery runbook must therefore explain how cluster-state rollback interacts with current application data.
k3s-killall.sh is not a normal upgrade command
Normal K3s service restarts deliberately allow Pod containers to continue running. K3s uses this behavior to reduce disruption during routine maintenance.
The k3s-killall.sh script is different.
K3s documents it as a way to:
- stop all K3s containers;
- reset containerd state;
- clean up K3s networking components and iptables chains;
- preserve cluster data.
The official rollback procedure uses k3s-killall.sh when all K3s processes must be stopped before datastore restoration.
K3s also warns that killall can cause data loss if applications are not shut down properly.
Therefore:
- do not use killall as the routine equivalent of
systemctl restart k3s; - drain or gracefully stop sensitive workloads when the recovery procedure allows;
- understand which applications use node-local or in-memory state;
- treat killall as a recovery/maintenance tool with a larger disruption boundary.
Upgrade validation should prove more than node version
A node reporting the new version proves only that the node changed.
A complete validation should cover four layers.
1. Control plane
- API responds normally;
- server nodes are Ready;
- controllers reconcile resources;
- etcd quorum or external datastore connectivity is healthy;
- no repeated control-plane errors appear in K3s logs.
2. Node platform
- expected K3s version is present;
- agents rejoin cleanly;
- CNI networking works across nodes;
- kubelet/runtime are healthy;
- cordoned nodes are intentionally uncordoned after validation.
3. Cluster services
- CoreDNS resolves Services;
- ingress accepts traffic;
- ServiceLB or external load-balancing paths remain healthy;
- storage provisioners/controllers run;
- monitoring and metrics remain available.
4. Applications
- critical Pods are Ready;
- StatefulSets mount the expected volumes;
- jobs and workers run;
- authentication works;
- real HTTP/API paths succeed;
- database writes and reads are correct;
- latency/error levels remain within the team's normal range.
A cluster upgrade is complete only when the workload platform is usable, not when every node has the same version string.
Production K3s upgrade runbook
Before maintenance
- Record the current K3s version on servers and agents.
- Select the exact target version or approved release channel.
- Read K3s release notes and version-specific upgrade caveats.
- Confirm no unsupported minor-version skip is planned.
- Review generic Kubernetes API/add-on compatibility separately.
- Create or verify the datastore recovery point.
- Back up
/var/lib/rancher/k3s/server/token. - Confirm PVC/database/application backup requirements.
- Verify how K3s configuration is persisted.
- Confirm server and agent upgrade order.
- Confirm HA server concurrency preserves quorum.
- Define pause, repair-forward, and rollback criteria.
- Reserve enough time for post-upgrade validation.
During maintenance
- Upgrade server nodes first.
- Change only one HA server at a time unless a tested topology explicitly supports more.
- Verify each server before changing the next.
- Verify the full control plane after the server set is complete.
- Upgrade agents afterwards.
- Watch System Upgrade Controller Plans/Jobs if automated.
- Stop the rollout immediately on an unexplained control-plane or datastore failure.
After maintenance
- Confirm every intended node reports the target version.
- Confirm all expected nodes are Ready and appropriately schedulable.
- Validate DNS, CNI, ingress, ServiceLB/external LB, and storage.
- Validate stateful applications and database access.
- Validate critical customer workflows.
- Review K3s and controller logs for recurring errors.
- Record the completed version and maintenance outcome.
- Keep the pre-upgrade recovery point according to the rollback window and retention policy.
Raff VM boundary for self-managed K3s upgrades
When K3s runs on Raff VMs, the cluster is self-managed.
That means the customer team owns:
- K3s version selection;
- release-note review;
- server and agent ordering;
- System Upgrade Controller configuration if used;
- host operating-system maintenance;
- datastore snapshots and server-token recovery;
- application-data protection;
- maintenance validation;
- K3s rollback and restore.
Raff provides the Linux VM and networking boundary, while K3s remains software your team operates on those VMs.
If the reason for adopting Kubernetes is the workload model rather than control-plane ownership, compare this responsibility set with Raff Kubernetes. A managed control plane changes who operates cluster lifecycle infrastructure, although workload API compatibility, application releases, data protection, and application validation still require customer ownership.