WebSocket hosting means running a server that can keep long-lived, two-way connections open with browsers, mobile clients, or other applications. For an early real-time product, one well-sized cloud VM is often the simplest place to start. The infrastructure becomes more complex when concurrent connections, message fan-out, availability requirements, or deployment risk force the application across multiple nodes.
For teams evaluating WebSocket hosting rather than only learning the protocol, the important questions are practical: how many connections must stay open, how much data moves through them, what happens when a server restarts, and how will the architecture scale without making the product unnecessarily expensive? Raff Technologies provides cloud VMs and supporting infrastructure that let teams run and control this stack directly, from a single WebSocket server to a multi-node deployment.
This guide explains those decisions without assuming that every real-time application needs Kubernetes, a managed real-time service, or a distributed architecture on day one.
What Is WebSocket Hosting?
A WebSocket server maintains persistent communication sessions instead of treating every interaction as a separate request-response cycle. The protocol begins with an HTTP handshake and then upgrades the connection so client and server can exchange messages in both directions over the same connection.
That makes WebSockets useful for workloads such as chat, collaborative applications, multiplayer features, live dashboards, notifications, and streaming application state. The protocol itself is standardized in RFC 6455, while the browser-facing API is documented by MDN.
The hosting difference is operational. A conventional web server may handle a large number of short requests. A WebSocket server can have thousands of clients connected while relatively few messages are moving. Capacity therefore depends on more than requests per second.
| Metric | Why it matters |
|---|---|
| Concurrent connections | Each live connection consumes some memory, socket/file-descriptor capacity, and network state |
| Messages per second | Drives application CPU and serialization work |
| Payload size | Affects bandwidth, memory, and serialization cost |
| Fan-out | One event sent to many clients can multiply work quickly |
| Connection duration | Long sessions change deployment and failure behavior |
| Reconnect rate | A restart or network problem can create a sudden connection spike |
| Message latency | Real-time products become visibly degraded when delivery slows |
A useful capacity model is therefore connections × activity per connection, not simply monthly visitors.
WebSockets vs SSE vs Polling
Do not choose WebSockets merely because the product is described as “real time.” Choose the simplest communication model that satisfies the interaction.
| Model | Good fit | Main trade-off |
|---|---|---|
| HTTP polling | Infrequent status checks and simple clients | Repeated requests add overhead when updates are frequent |
| Long polling | Compatibility-oriented applications with occasional updates | More request churn than a persistent connection |
| Server-Sent Events (SSE) | Primarily server-to-browser streams | Browser communication is one-way from server to client |
| WebSockets | Frequent bidirectional communication | Persistent connections require more operational planning |
If a dashboard only receives server updates, SSE may be sufficient. If a browser and server need to exchange frequent messages in both directions, WebSockets are usually the more natural model. If updates occur every few minutes, ordinary polling may be easier to operate.
This decision matters commercially too: unnecessary infrastructure is still infrastructure your team has to monitor, secure, deploy, and pay for.
When Should You Host a WebSocket Server on a Cloud VM?
A cloud VM is a strong fit when you want control over the WebSocket process, reverse proxy, runtime, network configuration, and scaling path without building around the restrictions of a platform that may impose connection-duration or protocol limitations.
Common fits include:
- SaaS applications adding chat, presence, live status, or collaboration.
- Internal dashboards with continuously changing data.
- Node.js, Go, Python, Java, Elixir, or other persistent application servers.
- Docker-based real-time services.
- Teams that want to start with one server and add infrastructure only when measured load requires it.
A VM is less attractive when the team explicitly wants a fully managed real-time messaging product and does not want to operate the connection layer. That can reduce operational work, although the application then needs to fit the provider's pricing, limits, APIs, and architecture.
The hosting decision is therefore not “VMs are always better.” It is control and predictable infrastructure ownership versus outsourcing more of the real-time layer.
How to Size a WebSocket Server
There is no reliable universal formula such as “1 GB RAM supports X WebSocket users.” Framework, runtime, TLS, authentication state, message buffers, application code, operating system limits, and traffic patterns all affect the result.
Start with four workload measurements:
- Peak concurrent connections. Measure simultaneous connected clients, not registered users.
- Average and peak message rate. Include inbound and outbound messages.
- Typical and maximum payload size. Small JSON events and large streamed payloads create very different network profiles.
- Fan-out ratio. A message delivered to one client is different from one broadcast to 10,000 clients.
Then load-test the actual application and observe CPU, memory, connection count, network throughput, event-loop or worker latency, and error rates. Keep headroom for traffic spikes and reconnect storms rather than planning to run permanently at the observed failure boundary.
For a first deployment, Raff Linux VMs provide a straightforward environment for running the application and measuring the real workload before introducing additional nodes.
Why Persistent Connections Change Memory and Network Planning
An idle WebSocket can consume little CPU, but it is not free. The operating system and application maintain socket state, and the application may retain authentication, subscription, room, or session information for each client.
Memory pressure often grows with connection count. CPU pressure often grows with message processing. Network pressure grows with message volume and fan-out. These limits can arrive at different times.
For example, two applications can both have 5,000 active connections:
- App A sends a small state update to each client once per minute.
- App B sends multiple updates per second and broadcasts many events to large groups.
The second workload can require substantially more CPU and network capacity even though the connection count is identical.
This is why buying a WebSocket server based only on vCPU count or a generic “number of users” estimate is risky. Benchmark the application pattern you actually expect.
Reverse Proxy Requirements for WebSocket Hosting
Production WebSocket servers commonly sit behind Nginx, Caddy, Traefik, or another proxy. The proxy can terminate TLS, route domains and paths, and separate public traffic handling from the application process.
The WebSocket handshake uses HTTP upgrade semantics. With Nginx, the Upgrade and Connection headers need explicit proxy handling because they are hop-by-hop headers; Nginx documents the current behavior in its WebSocket proxying documentation.
Check these settings when deploying:
| Proxy concern | What to verify |
|---|---|
| Protocol upgrade | WebSocket upgrade reaches the upstream correctly |
| TLS | Production connections use wss:// where appropriate |
| Idle timeout | Valid quiet connections are not terminated unexpectedly |
| Maximum connections | Proxy limits do not become the first bottleneck |
| Health checks | Unhealthy backends stop receiving new connections |
| Graceful shutdown | Existing connections can drain during maintenance when possible |
For the broader distinction between the two edge roles, see Reverse Proxy vs Load Balancer.
WebSocket Load Balancing: What Changes at Multiple Nodes?
WebSocket load balancing becomes relevant when one server is no longer enough for capacity or availability. The important difference from ordinary stateless HTTP is that a connected client remains attached to a backend for the lifetime of that connection.
A typical scaling path is:
| Stage | Architecture | Good fit |
|---|---|---|
| 1 | One VM running the WebSocket application | Prototype and early production |
| 2 | Reverse proxy + application on one VM | TLS, routing, cleaner process separation |
| 3 | Separate application/worker services | Background work begins affecting connection handling |
| 4 | Load balancer + multiple WebSocket nodes | Capacity or availability exceeds one node |
| 5 | Shared pub/sub or messaging layer | Events must reach clients connected to different nodes |
Adding a second WebSocket server does not automatically make the application distributed correctly. If user A is connected to node 1 and user B to node 2, an event created on node 1 may need a shared mechanism to reach node 2.
Depending on the application, that mechanism might be a pub/sub system, message broker, database-backed event flow, or another coordination layer. The correct choice depends on delivery requirements and workload; do not add one merely because a diagram of a WebSocket architecture contains it.
Do WebSockets Need Sticky Sessions?
Sticky sessions can keep a reconnecting or session-bound client associated with a particular backend, but they are not a substitute for distributed application design.
| Problem | Sticky sessions solve it? |
|---|---|
| Keep a connection on its established backend | The live connection already does this |
| Preserve backend-local state across new connections | Sometimes |
| Deliver an event to users on other nodes | No |
| Recover when a backend fails | No |
| Prevent uneven long-lived connection distribution | Not necessarily |
| Share presence or room membership globally | No |
Prefer to understand which state truly must be shared. A stateless reconnect path plus shared application state can be easier to recover than relying heavily on backend affinity.
Reliability: Design for Disconnects and Reconnect Storms
A production WebSocket application should assume that connections will disappear. Mobile network changes, browser suspension, proxy timeouts, deploys, VM restarts, backend failures, and ordinary internet instability can all disconnect clients.
Define the recovery behavior before traffic grows:
- Use reconnect backoff rather than allowing every client to retry continuously.
- Add jitter so thousands of clients do not reconnect at exactly the same interval.
- Decide whether a reconnect needs a full state refresh or event replay.
- Make duplicate event handling safe where delivery can be retried.
- Drain connections during planned deployments where the stack supports it.
- Test a server restart under realistic connection load.
A reconnect storm is an important capacity case. A server that comfortably maintains 20,000 established connections may behave very differently when thousands of clients perform authentication and connection setup simultaneously.
WebSocket Security Checklist
Persistent connections do not bypass normal application-security requirements. They create a long-lived public application surface.
Use a baseline that includes:
- TLS (
wss://) for production traffic where data crosses untrusted networks. - Authentication during or immediately after connection establishment.
- Authorization for channels, rooms, topics, and actions—not just authentication.
- Origin validation for browser clients where appropriate.
- Strict message validation and size limits.
- Rate limits for connection attempts and messages.
- Secret management outside application source code.
- OS and runtime patching.
- Logging for failed authentication, abnormal disconnects, and abuse patterns.
If internal services such as databases or message systems do not need public exposure, place them on private networking. Raff's VPC can be part of that design when separating public application nodes from private service traffic.
Observability for Real-Time Applications
CPU and RAM alone do not tell you whether a WebSocket product is healthy. Add connection-level signals.
Track at least:
| Signal | What it can reveal |
|---|---|
| Active connections | Capacity trend and abnormal drops |
| Connection attempts | Traffic spikes and abuse |
| Connection establishment failures | Proxy, TLS, auth, or capacity problems |
| Reconnect rate | Network, timeout, deployment, or backend instability |
| Messages per second | Application workload |
| Message processing latency | Event-loop or worker pressure |
| Outbound queue/buffer growth | Slow consumers or overloaded fan-out |
| Unexpected disconnects | Reliability regression |
| Network throughput | Transfer pressure |
| Pub/sub or broker lag | Cross-node delivery bottlenecks |
The browser WebSocket interface itself does not provide backpressure; MDN notes that applications can experience buffering problems when messages arrive faster than they can be processed. This is another reason to measure queue or buffer growth rather than watching CPU alone. MDN WebSocket API
For a broader monitoring model, see Application Observability for Small Teams.
How Much Does WebSocket Hosting Cost?
The advertised VM price is only one part of WebSocket hosting cost. A useful estimate includes the infrastructure needed to meet the application's actual reliability and scale target.
Model these components:
| Cost driver | When it grows |
|---|---|
| Compute and memory | More connections, heavier message processing, larger per-client state |
| Network transfer | Higher message frequency, payload size, and fan-out |
| Storage | Durable application state, logs, databases, and retained events |
| Load balancing | Multi-node traffic distribution becomes necessary |
| Shared messaging/state | Nodes must coordinate events or presence |
| Backup/recovery | Persistent application data needs recovery protection |
| Operations | Monitoring, incident response, deploys, and capacity testing require team time |
Raff uses monthly VM plans rather than hourly billing. Check the current Raff pricing page for current plan prices instead of relying on a hard-coded number in an architecture guide.
This distinction matters when comparing providers. A low entry price can be irrelevant if the workload later requires paid transfer, additional managed components, or a larger architecture than expected. Compare the total design you need, not only the smallest advertised server.
Single VM vs Multi-Node WebSocket Hosting
Start with one VM when it meets the requirement. Distributed systems create real operational cost.
One VM is usually enough when
- Peak connection count is comfortably within tested capacity.
- Message rate and fan-out leave CPU and network headroom.
- Short maintenance interruptions are acceptable.
- The team is still validating the product or traffic model.
- Cross-node state coordination would add complexity without solving a current problem.
Move toward multiple nodes when
- Load testing shows a real single-node capacity ceiling.
- One node is an unacceptable availability risk for the business.
- Deployments disconnect too many important sessions.
- A single process or VM cannot meet latency objectives at peak traffic.
- Traffic requires geographic or architectural separation.
When moving to multiple nodes, plan load balancing, health checks, connection draining, reconnect behavior, and shared state together. See Single Server vs Multi-Server Architecture before adding nodes only for appearance of scale.
A Practical WebSocket Hosting Architecture on Raff
A small team can use a staged architecture rather than buying the final architecture on day one.
Stage 1 — One cloud VM. Run the WebSocket application, reverse proxy, and monitoring on a Raff Cloud VM or Linux VM. Load-test the application and establish baseline connection, message, memory, CPU, and network metrics.
Stage 2 — Separate persistent data. As the application matures, move durable state to the appropriate database/storage layer and make sure application data has a recovery plan. Use Volumes where block storage fits the workload and Data Protection for supported backup/recovery needs.
Stage 3 — Isolate private services. Keep services that do not need public ingress on a VPC where the architecture benefits from private service communication.
Stage 4 — Add nodes for a measured reason. Introduce load balancing and a shared message/state mechanism only after capacity or availability requirements justify them.
This path keeps the initial hosting model understandable while leaving room to scale. It also makes provider evaluation easier: the team can see which infrastructure component solves which measured constraint.
