For the network side of the same HA design — Flannel, Traefik, ServiceLB, private node traffic, and the separate fixed registration/API endpoint — use K3s Networking, Ingress, and Load Balancing.
K3s does not use etcd by default on a normal single-server installation. If you do not configure another datastore and there is no embedded-etcd database already on disk, K3s uses embedded SQLite. SQLite cannot be used for a multi-server K3s cluster.
For the broader VM topology — single-server, agents, fixed registration endpoint, networking, storage, and recovery ownership — see K3s on Cloud VMs: Architecture, HA, Storage, and Operations.
Once you need a highly available K3s control plane, the datastore decision changes: use embedded etcd with three or more server nodes or use two or more K3s server nodes backed by a separately reliable external datastore such as PostgreSQL, MySQL, MariaDB, or etcd.
For most small teams building their first self-managed HA K3s cluster, embedded etcd is the cleaner starting point. An external database is stronger when database HA, backups, monitoring, and recovery already exist as a mature service or when larger-cluster requirements justify separating the datastore tier.
Raff Technologies can host the self-managed VM side of either architecture. If your team does not want to own datastore quorum, control-plane upgrades, and cluster recovery, compare that path with Raff Kubernetes before adding more self-managed infrastructure.
Single server: SQLite is the default. Multi-server HA: choose embedded etcd or an external datastore. Do not call an external-database design HA unless the database itself is also highly available.
K3s datastore options at a glance
| Topology | Datastore | What it means |
|---|---|---|
| Single K3s server | Embedded SQLite by default | Simplest model; not a multi-server HA datastore |
| HA with embedded etcd | 3+ K3s server nodes | Control-plane servers also form the etcd quorum |
| HA with external database | 2+ K3s server nodes + external DB | K3s servers share PostgreSQL, MySQL, MariaDB, or etcd |
The K3s project documentation currently states that SQLite is the default datastore when no other datastore is configured and no embedded-etcd database exists on disk. Embedded etcd is selected when K3s initializes or joins an etcd cluster, or when existing etcd data is detected.
That distinction matters because many searches for “K3s etcd” assume etcd is always the default. It is not.
Does K3s use SQLite or etcd by default?
For a fresh single-server K3s installation, SQLite is the default.
The practical progression is:
Single K3s server → embedded SQLite by default Need multiple K3s server nodes / control-plane HA → embedded etcd or → external datastore
SQLite is appropriate when:
- one server is enough;
- control-plane HA is not required;
- the cluster is development, edge, CI/CD, lab, or another workload where one-server simplicity is more valuable than redundant control-plane state.
SQLite cannot be the shared datastore for a cluster with multiple K3s server nodes.
If an existing SQLite-backed single-server cluster needs to become HA, K3s supports converting it to embedded etcd by restarting the server with --cluster-init, then joining additional server nodes.
Treat that conversion as a real infrastructure change: back up the datastore, preserve the server token, confirm the converted cluster is healthy, and add additional servers deliberately.
K3s HA has two datastore architectures
A highly available K3s control plane needs more than multiple server VMs. The datastore holding Kubernetes state must also remain available, consistent, and recoverable.
K3s offers two main multi-server HA patterns.
Embedded etcd
Stable registration endpoint ↓ ┌────────────┼────────────┐ ↓ ↓ ↓ Server 1 Server 2 Server 3 API + etcd API + etcd API + etcd └────────────┬────────────┘ ↓ Agents
Each K3s server participates in the control plane and embedded etcd cluster.
External datastore
Stable registration endpoint ↓ ┌────────┴────────┐ ↓ ↓ Server 1 Server 2 control plane control plane └────────┬────────┘ ↓ External datastore PostgreSQL / MySQL / etcd ↓ Agents
The K3s servers share a separate datastore rather than maintaining embedded-etcd membership themselves.
The real difference is not simply three servers versus two. It is where datastore availability, backup, performance, and recovery responsibility live.
Embedded etcd vs external database: quick decision
| Decision area | Embedded etcd | External datastore |
|---|---|---|
| Minimum K3s server nodes for HA | 3 | 2 |
| Datastore location | On K3s server nodes | Separate database tier/service |
| HA mechanism | etcd quorum | K3s server redundancy + database HA |
| External dependency count | Lower | Higher |
| Backup owner | K3s etcd snapshot workflow | Database-native backup system |
| Storage sensitivity | High on server nodes | Shifted to database tier |
| Server replacement | Must preserve healthy etcd membership | Less coupled to datastore membership |
| Best fit | Small/self-contained HA K3s | Mature DB platform or larger-cluster requirements |
Neither architecture is automatically safer merely because it has more components.
Choose the model whose failure and recovery paths your team can actually operate.
Embedded etcd requires quorum
K3s documentation requires an HA embedded-etcd cluster to have three or more server nodes and recommends an odd number of members for efficient quorum.
With three server nodes:
Server 1 ✓ Server 2 ✓ Server 3 ✕ 2 of 3 available → quorum remains
With only one of the three available:
Server 1 ✓ Server 2 ✕ Server 3 ✕ 1 of 3 available → quorum is lost
For a cluster with n members, etcd needs a majority. This is why adding a fourth member to a healthy three-member cluster does not increase tolerated failures: quorum rises to three while the cluster still tolerates only one failed member.
Operationally, avoid actions that can remove quorum at once:
- rebooting several server nodes together;
- upgrading all servers simultaneously;
- replacing multiple etcd members at once;
- changing critical firewall/network rules across all servers in one step;
- performing shared storage maintenance without considering the etcd set.
HA is partly an architecture property and partly a maintenance discipline.
Embedded etcd needs fast, reliable server storage
K3s performance depends heavily on datastore performance. The current K3s requirements recommend SSD-backed storage where possible and specifically note that etcd is write intensive.
That matters because an embedded-etcd server can be doing several jobs at once:
Kubernetes API + scheduler/controllers + etcd writes + system Pods + optionally user workloads + operating-system/runtime I/O
A server may have plenty of CPU available while slow disk latency still damages control-plane health.
For embedded etcd:
- prefer SSD/NVMe-backed server storage;
- monitor disk latency, not just capacity;
- keep control-plane headroom separate from workload sizing;
- avoid saturating server-node storage with unrelated write-heavy workloads.
K3s also requires embedded-etcd server nodes to communicate with each other on the etcd ports used for peer/client traffic. Keep that traffic on controlled private networking where possible rather than exposing etcd to the public internet.
External datastore HA is two HA systems, not one
K3s supports external etcd, PostgreSQL, MySQL, and MariaDB as cluster datastores.
For the external-database HA topology, K3s requires two or more server nodes that point at the same external datastore.
But this architecture:
K3s Server A ─┐ ├── Single PostgreSQL VM K3s Server B ─┘
is not fully highly available. The database is still a single point of failure.
A production external-datastore design must answer two separate questions:
- Can a K3s server fail without losing the control plane?
- Can the datastore fail without losing cluster state access?
The database tier therefore needs its own design for:
- instance failure;
- storage failure;
- failover;
- backups;
- restore;
- endpoint stability;
- monitoring;
- capacity;
- authentication;
- TLS;
- patching.
Moving the datastore outside K3s does not remove database operations. It moves them to another system.
External databases are strongest when they reuse existing operational maturity
An external datastore makes sense when your organization already operates a reliable database service or when cluster scale justifies a separate datastore tier.
Benefits can include:
- database storage tuned independently from K3s server VMs;
- centralized database backup and monitoring;
- server replacement without changing datastore membership;
- clearer separation between control-plane compute and persistent cluster state;
- reuse of existing database HA/recovery processes.
K3s project requirements currently recommend an HA setup with an external database for production and large clusters. That guidance does not mean embedded etcd is invalid for smaller production clusters; embedded etcd remains a supported HA architecture. The point is that larger environments can justify a separately managed datastore tier.
For a small team that would have to invent a new database platform solely to avoid a third K3s server, embedded etcd may be the simpler and more reliable system.
PostgreSQL and MySQL compatibility needs more than a connection string
K3s can use PostgreSQL, MySQL, or MariaDB through --datastore-endpoint, but the selected database service must support the behavior K3s expects.
One important current requirement is prepared statement support. The K3s datastore documentation notes that connection poolers such as PgBouncer may require additional configuration to work correctly with K3s.
That creates an important buying rule:
Do not assume every managed PostgreSQL endpoint is automatically suitable as a K3s datastore just because it accepts PostgreSQL connections.
Verify:
- direct database vs pooled endpoint behavior;
- prepared statement compatibility;
- TLS validation;
- connection limits;
- failover endpoint behavior;
- latency between K3s servers and the database;
- backup and restore process.
For Raff specifically, do not treat Managed PostgreSQL as a verified K3s datastore dependency unless the exact endpoint and compatibility path have been validated for that use case. A managed database can be excellent for application data without automatically being the right datastore for the Kubernetes control plane.
External etcd is different from embedded etcd
K3s can also use an external etcd cluster.
K3s Server A ─┐ K3s Server B ─┼── External etcd cluster K3s Server C ─┘
This separates etcd membership from K3s server lifecycle, but it does not eliminate etcd operations.
The team still owns:
- quorum;
- certificates;
- storage performance;
- backups;
- upgrades;
- restore testing;
- endpoint availability.
For most small teams, external etcd increases separation without necessarily reducing complexity. It fits better when etcd is already an intentionally operated platform.
Backup and HA protect against different failures
High availability keeps selected current failures from immediately stopping the control plane. Backups let you recover older state after destructive or logical failures.
| Failure | HA helps? | Backup helps? |
|---|---|---|
| One embedded-etcd server fails | Yes, while quorum remains | Usually not immediately required |
| One K3s server fails with external DB | Yes, if another server and DB remain healthy | Usually not immediately required |
| Accidental object deletion | No | Yes |
| Bad configuration replicated across the datastore | No | Yes |
| Datastore corruption | Maybe not | Yes |
| Entire control-plane environment lost | Depends on architecture | Yes |
| Required restore token lost | No | Only if token was preserved |
Do not use three etcd members as a substitute for backups.
Do not use backups as a substitute for control-plane HA when downtime is unacceptable.
Embedded-etcd snapshots: know the defaults, then set your own policy
K3s provides native k3s etcd-snapshot tooling for embedded etcd.
Current K3s documentation says scheduled etcd snapshots are enabled by default:
- at 00:00 and 12:00 system time;
- with five scheduled snapshots retained by default.
Snapshots can also be created on demand and optionally uploaded to S3-compatible object storage.
Those defaults are a starting point, not an RPO/RTO strategy.
Define your own:
- snapshot frequency;
- retention;
- off-node/off-site copy;
- encryption requirements;
- restore test schedule;
- ownership for failed snapshot jobs.
If using an S3-compatible target, verify the exact object-storage compatibility and retention controls required by your recovery policy.
External datastore backups happen outside K3s
When K3s uses an external datastore, K3s does not replace the database's backup and restore system.
The database owner must define:
- recurring backups or snapshots;
- retention;
- point-in-time recovery where supported;
- replica/failover behavior;
- restore testing;
- recovery credentials;
- monitoring for backup failures.
A database that is “managed” but has an untested restore path is still a weak control-plane dependency.
Preserve the K3s server token
This is an easy recovery detail to miss.
K3s documentation requires you to back up the server token at:
/var/lib/rancher/k3s/server/token
The token is used to encrypt confidential data stored in the datastore. When restoring, the same token value is required; a datastore snapshot without the required token can be unusable.
Backup policy should therefore treat the token as part of the cluster recovery set, not as an incidental file.
Stable registration address is separate from datastore HA
An HA cluster also needs a stable way for agents and administrators to reach the K3s server layer.
K3s documentation suggests approaches such as:
- a Layer 4 TCP load balancer;
- round-robin DNS;
- a virtual or elastic IP.
This stable registration/API endpoint solves a different problem from datastore HA.
You need both:
Stable control-plane endpoint ↓ Multiple K3s servers ↓ Available datastore
A healthy etcd cluster does not help an agent if its only configured server endpoint disappeared. Likewise, a perfect registration endpoint cannot compensate for a failed datastore.
Cost: compare the complete architecture
External DB can reduce the K3s server requirement from three to two, but that does not automatically make it cheaper.
Compare the full system:
Embedded etcd 3+ K3s server VMs + snapshot storage + quorum-aware maintenance External datastore 2+ K3s server VMs + HA database service/infrastructure + database backups + monitoring + database lifecycle + network dependency
If the external database platform already exists, its incremental cost may be small.
If you must create a new HA database solely for K3s state, the architecture can cost more—in both infrastructure and operator time—than a three-server embedded-etcd cluster.
Decision framework for a small team
Choose SQLite when:
- one K3s server is enough;
- control-plane HA is not required;
- simplicity is more important than redundant control-plane state.
Choose embedded etcd when:
- you need multi-server HA;
- three K3s server VMs are acceptable;
- you want a self-contained control plane;
- SSD-backed server storage is available;
- the team can operate quorum and snapshot recovery;
- you do not already have a stronger external database platform.
Choose an external datastore when:
- a reliable HA PostgreSQL/MySQL/MariaDB/etcd service already exists;
- backup and failover are mature and tested;
- separating datastore performance from K3s servers solves a real requirement;
- larger cluster scale justifies the extra tier;
- endpoint compatibility and prepared-statement behavior have been verified.
Avoid an external datastore if it would just be one ordinary database VM behind two K3s servers.
Avoid embedded etcd if the team cannot maintain quorum, fast server storage, and tested snapshot recovery.
Raff VM vs Raff Kubernetes for this decision
Self-managed K3s can run on Raff VMs when your team wants direct ownership of the servers and datastore architecture.
An embedded-etcd layout can look like:
3 K3s server VMs ↓ embedded etcd quorum ↓ agent VMs as workload capacity grows
An external-datastore layout can look like:
2+ K3s server VMs ↓ verified HA external datastore ↓ agent VMs
Use Raff VPC for private server-to-server traffic where appropriate. Keep the K3s API, etcd, database, and node network boundaries intentional rather than exposing internal control-plane dependencies publicly.
Raff VMs remain self-managed at the guest OS and application/platform level. Your team owns K3s installation, server patching, datastore maintenance, upgrades, snapshots, restore testing, and cluster recovery.
If that operational ownership is the part you are trying to avoid, Raff Kubernetes is the more relevant comparison. It changes the decision from “which datastore should we operate?” to “do we want to operate the Kubernetes control plane ourselves at all?”