Server health checks should answer one operational question and trigger one appropriate action. A liveness check asks whether the process should be restarted. A readiness check asks whether the instance should receive traffic. A startup check protects slow initialization. Synthetic monitoring asks whether users can actually reach and use the service.
That distinction matters because a bad health check can make an outage worse. Restarting every application instance during a database problem, routing users to an app that is still warming up, or declaring a service healthy because /health returns 200 while login is broken are all common failure modes.
For teams running applications on Raff Technologies, health checks are part of the production design, not a monitoring afterthought. Whether the workload runs on a single cloud VM, several VMs, Docker, or Kubernetes, the goal is the same: detect the right failure, take the least destructive recovery action, and give the team enough context to understand what happened.
Server health checks at a glance
| Check | Question | Typical action |
|---|---|---|
| Liveness | Is the process healthy enough to keep running? | Restart only when restart is likely to help |
| Readiness | Can this instance safely receive production traffic? | Remove from traffic until ready |
| Startup | Has initialization completed? | Delay liveness/readiness enforcement |
| Dependency check | Is a required downstream dependency usable? | Degrade, stop traffic, or alert depending on impact |
| Synthetic check | Can a user-facing workflow succeed from outside? | Alert/investigate customer impact |
| Worker/job health | Is background work actually progressing? | Alert or restart the affected worker if appropriate |
The key rule is simple: do not use one endpoint to control restarts, routing, uptime alerts, and business health at the same time.
Liveness vs readiness: the most important distinction
A service can be alive but not ready.
For example, a web process can still respond to a local health endpoint while its database pool is exhausted. Restarting the process may not fix the dependency problem, but continuing to send production traffic to it may still be harmful.
That leads to two different decisions:
- Liveness: should this process continue running?
- Readiness: should this instance receive new traffic right now?
Kubernetes documents the same separation: liveness probes determine when a container should be restarted, while readiness probes determine whether it should receive traffic. Incorrect liveness design can create cascading failures by restarting containers that are otherwise recoverable. Kubernetes probes
A practical rule is:
Liveness should fail only when restarting this process is likely to improve the situation. Readiness should fail when the instance should stop receiving new traffic.
What should a liveness probe check?
A good liveness probe is narrow, fast, and local to the process.
Useful liveness signals can include:
- the process can respond;
- the event loop or worker is not deadlocked;
- the runtime can make basic progress;
- critical internal state is not irrecoverably corrupted.
Avoid putting every dependency into liveness.
| Liveness check | Good idea? | Why |
|---|---|---|
| Process/event-loop can respond | Yes | Restart may recover a stuck process |
| Internal deadlock detector | Yes | Restart may restore progress |
| Full database query | Usually no | Restarting the app does not fix a DB outage |
| Payment provider call | No | External provider failure should not restart your app fleet |
| Object storage check | Usually no | Feature impact may be partial rather than process failure |
| Expensive business workflow | No | Too deep and fragile for restart control |
If a database outage makes every liveness probe fail, an orchestrator can restart the whole application fleet at once. That adds connection churn and startup load to an existing outage.
What should a readiness probe check?
A readiness probe should represent whether the instance can serve useful production traffic now.
That may include:
- application initialization completed;
- configuration loaded;
- required local components are ready;
- essential database connectivity is available;
- the instance is not draining during deployment;
- a critical dependency required by nearly every request is usable.
Readiness can be deeper than liveness because the action is less destructive: stop routing new traffic to the instance while it recovers.
Readiness should follow the traffic model
Do not automatically mark the entire service unready because one optional feature is degraded.
For example:
- If email delivery fails but checkout still works, the web app may remain ready while the email subsystem alerts separately.
- If the primary database is unavailable and almost every request requires it, readiness may reasonably fail.
- If image processing workers are down but the main API can still serve reads, the API may remain ready while the worker health signal fails.
The right boundary depends on what the instance is responsible for serving.
Health check endpoint design
A health check endpoint should be easy for humans and automation to interpret. Keep response behavior deterministic and inexpensive.
A basic split might look like:
/health/live -> process survival /health/ready -> traffic readiness /health/start -> startup completion, if your platform uses a separate check
The endpoint response should not expose secrets, credentials, detailed stack traces, or internal topology to public callers.
Good properties include:
- small response body;
- fast timeout;
- predictable status codes;
- no expensive queries;
- no unnecessary dependency chain;
- clear ownership of what failure means.
If an endpoint checks five external services and takes several seconds, it is no longer a simple routing signal. It is closer to a deep diagnostic or synthetic test.
Startup probes prevent restart loops
Some applications need meaningful startup time. JVM or .NET services may warm up; applications may load models or indexes; caches may initialize; a service may need to establish network connections before it can serve traffic.
A startup probe gives the application a realistic initialization window before normal liveness/readiness rules take effect. Kubernetes specifically uses startup probes to delay liveness and readiness until startup succeeds. Kubernetes startup probes
Use a startup check when:
- boot time is consistently longer than normal liveness thresholds;
- cold start behavior varies significantly;
- initialization includes large data or model loading;
- startup frequently triggers false liveness failures.
Do not use a huge startup window to hide an unknown boot problem. Measure normal startup duration and choose thresholds deliberately.
Dependency health checks: which dependencies belong where?
Dependencies should be checked according to the action you want to take when they fail.
| Dependency | Better use |
|---|---|
| Primary database | Often readiness if most traffic requires it |
| Cache | Readiness only if service cannot degrade without it |
| Queue/broker | Worker health or feature-specific alert |
| Payment provider | Synthetic or transaction-specific monitoring |
| Email provider | Feature-specific alert, rarely liveness |
| Object storage | Readiness only for services that cannot serve without it |
| Auth provider | Readiness or synthetic login check depending on architecture |
A dependency failure is not automatically a process failure.
The useful question is: what should the platform do if this dependency is unavailable? If the answer is “restart the app,” it may belong in liveness. If the answer is “stop sending this instance new traffic,” it may belong in readiness. If the answer is “alert the team but keep serving other features,” it belongs elsewhere.
Synthetic monitoring: the outside-in health check
Internal health endpoints cannot prove the whole customer path works.
Synthetic monitoring tests the service from outside, closer to the way a user experiences it. It can verify:
- DNS resolution;
- TLS and certificate validity;
- public HTTP availability;
- login;
- API responses;
- checkout or another critical transaction;
- file upload/download;
- WebSocket connection and basic message flow.
Google SRE distinguishes white-box monitoring, which uses internal system knowledge, from black-box monitoring, which evaluates externally visible behavior. Black-box monitoring is especially useful for active user-facing symptoms. Google SRE — Monitoring Distributed Systems
Uptime monitoring vs synthetic monitoring
Simple uptime monitoring asks whether an endpoint responds. Synthetic monitoring can go further by testing a meaningful path.
| Test | What it proves |
|---|---|
| Ping/TCP check | Network endpoint is reachable |
HTTP 200 check | A page/endpoint responds |
| API synthetic | Authenticated request returns expected result |
| Browser synthetic | User-facing flow can complete |
| Transaction synthetic | Critical business workflow still works |
A homepage returning 200 does not prove login works. A load balancer with healthy targets does not prove customers can complete checkout. Use synthetic tests for the small number of journeys that represent real service usability.
False positives and false negatives
A health-check system can fail in two directions.
False positive: the system is treated as unhealthy when it is actually usable.
False negative: the system is treated as healthy while customers are affected.
Examples:
| Failure | Type | Result |
|---|---|---|
| Liveness depends on a third-party API | False positive | Healthy app processes restart during provider outage |
| Readiness fails on optional cache | False positive | Capacity is removed unnecessarily |
/health always returns 200 | False negative | Broken business path stays “healthy” |
| Web worker responds but queue is stuck | False negative | Background work silently stops |
| Server is healthy but DNS fails | False negative | Users cannot reach it |
The solution is not “deeper checks everywhere.” It is to give each check a narrow job and combine internal health with black-box validation.
Health checks for load balancers and reverse proxies
Traffic-distribution systems should use readiness, not simple process existence, to decide whether a backend receives new requests.
A useful lifecycle is:
Starting -> not ready Initialized -> ready Draining -> not ready, still alive Broken process -> liveness failure Recovered -> ready again
This supports safer deployments because a server can stop receiving new traffic before the process is terminated.
See Reverse Proxy vs Load Balancer and Load Balancing Explained for the surrounding traffic architecture.
Health checks during deployment
Deployments are where badly designed probes often become visible.
A safer deployment sequence is:
- Start the new instance/process.
- Let startup complete.
- Wait for readiness.
- Add it to traffic.
- Mark the old instance unready/draining.
- Let active requests finish where practical.
- Stop the old process.
This prevents traffic from reaching a server simply because the process exists.
For multi-node applications, readiness should be part of the release strategy rather than a separate monitoring feature.
Background workers need different health checks
A background worker can be “alive” while doing no useful work.
For workers, add progress-oriented signals such as:
- last successful job time;
- queue depth;
- age of the oldest job;
- processing latency;
- failure/retry rate;
- heartbeat from the worker process.
A worker health endpoint that only returns 200 can miss the most important failure: jobs are not moving.
The same principle applies to cron jobs and scheduled tasks. Their health signal is often “did this job finish when expected?” rather than “does a process respond on a port?”
Database health is more than connection success
For databases, a single successful connection is only one signal.
Depending on the workload, important health signals can include:
- connection success;
- query latency;
- replication lag;
- disk capacity;
- lock/contention pressure;
- backup success;
- restore readiness.
Do not put all of those into your web application's liveness endpoint. Monitor the database as its own production component and let application readiness reflect only the dependency behavior needed for serving requests.
Alerting: do not page on every failed probe
A failed health check is a signal, not automatically a page-worthy incident.
Google SRE recommends separating signals that require urgent human action from lower-priority information. Monitoring should help answer both what is broken and why, while customer-facing symptoms should strongly influence paging decisions. Google SRE — Monitoring Distributed Systems
A practical model:
| Event | Typical urgency |
|---|---|
| One readiness failure during deploy | Usually no page |
| One redundant instance restarts once | Watch/ticket depending on context |
| Many instances fail readiness | High urgency |
| Synthetic login fails from several locations | High urgency |
| Startup never completes after realistic window | Deployment investigation |
| Optional dependency degraded | Warning/ticket |
| Public endpoint unavailable | Page when customer-impacting |
Alert on sustained, actionable symptoms rather than every transient probe result.
Health checks and observability are different
A health check answers a small operational question. Observability explains behavior.
If readiness fails, you still need metrics, logs, and traces to understand why. Health checks should therefore work alongside:
- CPU and memory metrics;
- latency and error rates;
- logs;
- traces;
- database metrics;
- queue metrics;
- dependency telemetry.
Google's SRE framework highlights latency, traffic, errors, and saturation as core monitoring signals for user-facing systems. Health checks complement these signals; they do not replace them.
See Application Observability for Small Teams for the broader monitoring model.
Health checks on Raff infrastructure
Raff gives teams control over the application and VM layers needed to implement custom health logic.
A practical small-team path is:
| Stage | Health design |
|---|---|
| One application VM | /live + /ready + external synthetic check |
| Reverse proxy in front | Route only to ready app processes |
| Multiple VMs | Per-node readiness + external synthetic + shared observability |
| Worker tier | Worker heartbeat + queue progress |
| Private dependencies | Keep database/internal-service traffic on VPC where appropriate |
| Business-critical workload | Add tested backup/recovery through Data Protection and an incident runbook |
For a simple web application, start on a Raff Cloud VM or Linux VM with clear application-level health endpoints. If the architecture later moves to Managed Kubernetes, keep the same semantic distinction between startup, liveness, and readiness instead of inventing new meanings for each environment.