Application observability helps a small team answer four production questions: Is the service working for users, what changed, where is the failure, and what should we do next? The practical starting point is not a large monitoring platform. Start with user-facing latency, traffic, errors, and saturation; add structured logs and reliable request IDs; then introduce distributed tracing when requests cross enough services that logs no longer explain the path.
For teams running workloads on Raff Technologies, the same rule applies: keep the first observability layer small enough to operate, and separate telemetry components only when production risk, data volume, or resource contention justifies another VM or service. Observability should shorten detection and recovery time—not become another platform the team struggles to maintain.
Small-team default: external availability checks + application metrics + infrastructure metrics + structured logs. Add tracing after architecture complexity creates a real debugging problem.
Application observability vs monitoring: the quick answer
Monitoring and observability overlap, but they are not identical.
Monitoring checks conditions you already know to watch: availability, error rate, disk space, queue depth, certificate expiry, backup failure, or CPU pressure.
Observability gives you enough telemetry and context to investigate why the system is behaving the way it is, including failures you did not predict in advance.
| Question | Monitoring | Observability |
|---|---|---|
| Is the service available? | Strong | Strong |
| Is latency above a threshold? | Strong | Strong |
| Which deployment changed behavior? | Limited without context | Strong |
| Which dependency caused one slow request? | Limited | Strong with traces/context |
| Why did only one workflow fail? | Often incomplete | Stronger with correlated signals |
| Who should act now? | Alert/runbook concern | Alert + diagnostic context |
The best small-team system uses both: monitoring should tell you what matters now, and observability should help explain why.
Metrics, logs, and traces answer different questions
OpenTelemetry treats metrics, logs, and traces as separate telemetry signals. They are most useful when they share consistent service, environment, deployment, request, and trace context.
| Signal | Best question | Practical use |
|---|---|---|
| Metrics | Is behavior changing? | Detection, trends, capacity, SLIs |
| Logs | What happened inside this component? | Errors, events, jobs, deployment context |
| Traces | Where did this request spend time or fail? | Cross-service latency and dependency diagnosis |
A latency metric can show that checkout became slower. Logs may reveal database timeouts. A trace can show that most of the request duration came from one external API.
Do not collect all three signals at maximum detail by default. Collect the minimum detail that supports detection, diagnosis, and recovery.
Start with the four user-facing service signals
Google SRE's four golden signals remain a useful baseline for user-facing systems:
- Latency: how long successful and failed requests or jobs take.
- Traffic: how much demand the system is serving.
- Errors: how often users receive incorrect or failed outcomes.
- Saturation: how close a constrained resource is to its practical limit.
A small web application might begin with:
| Signal | Example |
|---|---|
| Latency | p50, p95, p99 HTTP response time |
| Traffic | requests/second or completed jobs/minute |
| Errors | 5xx rate, failed jobs, rejected transactions |
| Saturation | CPU, memory pressure, DB connections, queue depth |
Do not use averages alone. A healthy average can hide tail latency, short saturation periods, or failures isolated to one endpoint.
The useful sequence is:
User symptom ↓ Service metric ↓ Application / dependency evidence ↓ Infrastructure evidence ↓ Action
That ordering prevents teams from treating every CPU spike as a customer incident.
Infrastructure metrics explain capacity, not user experience by themselves
For each production VM, database, cache, queue, and worker, collect the signals that reveal resource pressure and waiting.
A practical baseline includes:
- CPU utilization and load;
- memory use and swap activity;
- free disk space and growth rate;
- disk latency and I/O wait;
- network throughput and errors;
- process health and restart count;
- database connections and query latency;
- queue depth, retries, and oldest-job age;
- backup age and failure where recovery depends on it.
Interpret these beside application behavior. High CPU is not automatically an incident if latency and errors remain healthy. Low CPU does not prove health if the workload is waiting on storage, locks, or an external dependency.
For capacity diagnosis, use Cloud Server Performance Bottlenecks and Choosing the Right VM Size.
Structured logs should preserve searchable context
Logs become operationally valuable when the team can filter and correlate them during an incident.
Prefer structured records with consistent fields such as:
- timestamp;
- severity;
- service and component;
- environment;
- deployment/version;
- request ID or job ID;
- trace ID when tracing is enabled;
- route or workflow name;
- result/status code;
- duration;
- dependency name;
- retry count.
Standardize field names across services. If one service calls the field service, another calls it app_name, and a third calls it component, incident searches become slower and dashboards become harder to reuse.
Do not log passwords, private keys, access tokens, session cookies, payment data, or unnecessary personal data. Logs are copied, indexed, retained, and accessed by more systems than application developers often expect.
Application logs and audit logs also serve different purposes. Application logs explain software behavior; audit logs preserve accountable actions. See Application Logs vs Audit Logs.
Correlation IDs are a high-return improvement
Before a small team invests heavily in distributed tracing, consistent request and job IDs can provide much of the investigative value.
A practical flow is:
Incoming request ↓ request_id API / web application ↓ same request_id Worker / queue / downstream call ↓ same request_id Searchable logs
When an alert reports elevated errors, an operator can move from the metric to one failed request and then search every relevant component using the same ID.
Generate or accept an identifier at the edge, propagate it across important internal boundaries, and include it in logs. When OpenTelemetry tracing is later introduced, preserve trace and span context as well.
Add distributed tracing when request paths become hard to explain
Distributed tracing becomes valuable when a user action crosses several services, queues, workers, databases, or external APIs.
Add tracing when:
- metrics prove there is a latency problem but not where it occurs;
- one workflow crosses multiple services or VMs;
- asynchronous work hides the original failure;
- retries and fallbacks obscure the first error;
- different teams own different parts of one request;
- p95 or p99 latency is difficult to attribute.
A monolith on one VM may not need full distributed tracing. A multi-service application may need it early. Architecture complexity—not fashion—should determine the timing.
Trace volume can grow quickly, so sampling should be deliberate. Preserve enough healthy traffic to understand normal behavior while prioritizing errors, slow requests, low-volume critical workflows, and newly changed services when the backend supports those policies.
OpenTelemetry gives small teams a portable collection layer
OpenTelemetry is a vendor-neutral framework for generating, collecting, and exporting telemetry. The OpenTelemetry Collector can receive traces, metrics, and logs, process them, and forward them to one or more observability backends.
That separation can be useful because instrumentation does not have to be tightly coupled to one backend.
A simple architecture is:
Application / VM telemetry ↓ OpenTelemetry Collector ↓ Metrics / logs / traces backend
The Collector does not remove the need to choose a backend, retention policy, alerting model, or capacity plan. It provides a consistent collection and processing layer.
For a very small application, adding a Collector before it solves a real problem may still be unnecessary. Use it when standardizing telemetry across services or keeping backend options flexible creates operational value.
Three observability stages for small teams
Stage 1 — One production VM
Use the smallest reliable set:
- external uptime/black-box check;
- application error reporting;
- latency, traffic, and error metrics;
- VM CPU, memory, disk, and network metrics;
- structured application logs;
- deployment markers;
- backup-failure alerts.
At this stage, correlation IDs often provide more value than full tracing.
Stage 2 — Application and data services are separated
Once the app, worker, database, cache, or queue live on separate resources, add:
- dependency latency/error metrics;
- central log collection;
- consistent service/environment labels;
- request and job context across components;
- database and queue saturation indicators;
- an OpenTelemetry Collector if it reduces agent/instrumentation fragmentation.
Stage 3 — Multi-service or distributed application
When cross-service diagnosis becomes expensive, add:
- distributed tracing for important workflows;
- trace-to-log correlation;
- sampling and retention policies;
- service-level indicators and objectives;
- deployment-aware dashboards;
- stronger alert routing and runbooks.
The maturity target is not “collect everything.” It is “make the next incident faster to understand.”
Define SLIs before creating more alerts
A service level indicator (SLI) measures behavior users care about. A service level objective (SLO) sets a target for that indicator over a period.
Small teams can start with one or two critical workflows:
| Workflow | Possible SLI |
|---|---|
| Public API | successful valid requests under a latency threshold |
| Checkout | completed transactions / valid attempts |
| Authentication | successful valid logins |
| Background export | jobs completed within the expected window |
This prevents alerting from becoming a list of infrastructure thresholds with no connection to user impact.
An SLO does not mean every miss should page someone. It gives the team a shared definition of reliability and a stronger basis for deciding which symptoms require urgent action.
Alerts should be urgent, important, actionable, and owned
Prometheus guidance recommends keeping alerting simple and paging on symptoms associated with user pain rather than every possible cause.
For every urgent alert, define:
- the user-facing condition that is failing;
- the affected service/workflow;
- severity;
- an owner or rotation;
- the first diagnostic link or query;
- the first response step;
- how recovery is confirmed.
Use different destinations for different urgency levels:
| Signal | Better destination |
|---|---|
| Immediate user impact requiring action | Page / urgent incident channel |
| Capacity risk requiring planned work | Ticket / operations queue |
| Diagnostic cause without current impact | Dashboard / non-urgent notification |
| Deployment/configuration event | Timeline / searchable logs |
Avoid duplicate paging at every layer. If an application latency alert already captures user impact, separate urgent pages for every downstream CPU, disk, and database symptom may create noise instead of clarity.
Deployment markers reduce diagnosis time
Many incidents follow a change. Record deployment, configuration, migration, and infrastructure events on the same timeline as service metrics.
Preserve at least:
- deployment version;
- start and completion time;
- environment;
- actor or automation identity;
- result;
- rollback event when used.
During an incident, “what changed before the graph moved?” should not require reconstructing a timeline from CI logs, chat messages, and shell history.
Control observability cost before volume becomes a problem
Observability cost grows through several independent dimensions:
- log bytes ingested;
- trace/span volume;
- metric cardinality;
- retention duration;
- indexing and query cost;
- network transfer to an external backend;
- compute and storage for a self-hosted stack.
The first controls should be simple:
- Remove debug logs that are never used.
- Keep metric labels bounded; avoid raw user IDs, request IDs, or unbounded URL values as metric labels.
- Use route templates such as
/users/:idrather than unique paths. - Retain high-resolution data only as long as it is operationally useful.
- Sample traces deliberately.
- Keep request-specific detail in logs/traces rather than exploding metric cardinality.
Do not optimize away evidence that your incident process actually needs. Cost control should remove low-value telemetry, not the signals that prove user impact.
Do not put every monitoring component on the same failure boundary
If the application, metrics backend, logs, dashboards, and alerting all depend on one VM, that VM can fail and remove the evidence needed to diagnose itself.
For a small workload, co-locating components may still be a reasonable cost decision. Understand the trade-off and keep at least one external signal such as uptime/black-box monitoring outside the primary failure boundary.
As the service becomes more important, separate the monitoring backend when:
- telemetry competes with the application for CPU, memory, or disk;
- losing one VM would remove both the service and its evidence;
- log/metric retention requires materially more storage;
- several VMs need one shared observability system;
- incident diagnosis depends on telemetry surviving application-host failure.
This is also a reason to monitor the monitoring system itself.
Self-hosted vs managed observability is an ownership decision
A self-hosted stack can offer control over software, retention, and deployment. A managed observability service can reduce the work required to operate storage, upgrades, alerting components, and backend availability.
Choose based on operational ownership rather than ideology.
Self-hosting is a stronger fit when:
- the team is comfortable operating the stack;
- data volume is predictable;
- storage and retention are understood;
- control over deployment and data location matters;
- the team can recover the observability backend after failure.
A managed backend is often a stronger fit when:
- the team wants to spend less engineering time maintaining telemetry infrastructure;
- retention or ingestion scales unpredictably;
- backend availability matters more than infrastructure control;
- the organization needs capabilities the team does not want to operate itself.
The collection layer and backend decision can also be separated. OpenTelemetry can provide portable instrumentation while the backend remains managed or self-hosted.
How this applies on Raff
A Raff VM can host an application, a collector, or a small self-hosted observability stack. As telemetry volume or architecture complexity grows, move the observability backend to a separate VM instead of allowing monitoring workloads to compete with the application they are supposed to diagnose.
A practical Raff path is:
Stage 1 Raff VM: application + lightweight telemetry agents External uptime check Stage 2 Raff VM(s): application / workers ↓ private traffic Raff VPC ↓ Dedicated telemetry VM or external backend Stage 3 Application VMs + database + workers ↓ Collector / observability backend ↓ Separate retention and recovery policy
Use Raff VPC for private service-to-service communication where supported. If a self-hosted telemetry backend needs storage growth independent from the application VM, evaluate Raff Volumes as an infrastructure storage layer and design the guest filesystem and application retention policy explicitly.
Raff Data Protection can support infrastructure-level recovery for eligible resources, but backups do not replace application-aware export, retention, or restore procedures for the observability software itself.
Raff provides the infrastructure boundary. Your team still owns instrumentation, alert rules, telemetry retention, access control, dashboards, incident workflows, and restore testing unless a managed service explicitly transfers that responsibility.