Kubernetes monitoring is the process of collecting and interpreting metrics, logs, Events, and traces so a team can detect failures, understand what changed, and decide what to do next.
For small teams, the goal is not maximum telemetry. It is to collect the smallest set of signals that answers production questions reliably.
The practical model is:
metrics → what changed
Events → what Kubernetes reported while reconciling
logs → what the application or component said
traces → where one request spent time
Raff Technologies provides a built-in infrastructure baseline for managed Kubernetes—cluster/node CPU and memory, workload views, live Pod logs, and an activity feed—while application-level service metrics, alert policy, durable log retention, and distributed tracing remain workload responsibilities.
This guide owns the monitoring and observability model. It does not prescribe one mandatory Prometheus, Grafana, Loki, or tracing stack.
Monitoring and observability are related, but not identical
Monitoring asks whether known conditions are healthy.
Examples:
- Is error rate above the acceptable threshold?
- Are replicas unavailable?
- Is a node under memory pressure?
- Are Pods restarting?
- Has queue depth exceeded the operating range?
Observability is broader. Kubernetes describes observability through the collection and analysis of metrics, logs, and traces to understand the internal state, performance, and health of a cluster.
Kubernetes Events are useful alongside those three pillars, but they serve a different role: they are short-lived diagnostic records about Kubernetes activity rather than a durable observability store.
| Signal | Best question | Example |
|---|---|---|
| Metrics | What changed, by how much, and for how long? | CPU, memory, request rate, p95 latency, error rate |
| Events | What did Kubernetes report while reconciling state? | FailedScheduling, image pull failure, eviction |
| Logs | What did this process say happened? | stack trace, timeout, auth failure |
| Traces | Where did this request spend time? | API → worker → database latency |
A production incident should start from the signal closest to user or workload impact, then move down the stack.
Metrics should be split into infrastructure, Kubernetes, and application layers
One metrics dashboard cannot answer every operational question.
A useful small-team model has three layers.
Infrastructure metrics
Examples:
- node CPU;
- node memory;
- disk pressure;
- node availability;
- network saturation where relevant.
These answer:
Is the worker infrastructure healthy enough to run workloads?
Kubernetes workload metrics
Examples:
- desired vs available replicas;
- Pod restart counts;
- Pending Pods;
- failed Jobs;
- rollout status;
- HPA replica changes.
These answer:
Is Kubernetes maintaining the desired workload state?
Application metrics
Examples:
- request rate;
- error rate;
- latency;
- queue depth;
- job completion rate;
- checkout failures;
- webhook failures.
These answer:
Is the application delivering the intended service?
The third layer is the one small teams most often miss.
A Deployment can show all replicas Ready while every request returns HTTP 500.
Infrastructure health is not the same thing as application health.
Metrics Server is not a complete monitoring system
Kubernetes resource metrics are intentionally narrow.
The Metrics API provides current CPU and memory usage for Nodes and Pods. It supports tools such as:
kubectl top nodes kubectl top pods
and resource-metric-based autoscaling.
Kubernetes 1.37 promoted the Metrics API to stable as metrics.k8s.io/v1.
That graduation improves the API's stability guarantees. It does not turn the resource metrics pipeline into a historical monitoring platform.
Metrics Server still does not replace:
- historical time-series storage;
- service-level metrics;
- dashboards over long periods;
- alerting;
- custom business metrics;
- durable trend analysis.
A production metrics pipeline normally needs a scraper and time-series backend when historical visibility or alerting is required.
Prometheus is a common choice, but not a mandatory one.
Use the system that answers the operating questions your team actually has.
Kubernetes component metrics provide a deeper cluster view
Kubernetes components expose operational metrics that can be scraped by monitoring systems.
Common sources include:
- kube-apiserver;
- kube-scheduler;
- kube-controller-manager;
- kubelet;
- kube-proxy.
The kubelet also exposes specialized metrics endpoints such as resource, cAdvisor, and probe metrics.
These signals are useful for questions such as:
- Is API-server latency increasing?
- Are scheduler operations failing?
- Are admission operations becoming slow?
- Are node-level workload metrics changing?
They are valuable when operating the cluster, but they still do not replace application-level indicators.
If users report an API failure, the correct recovery signal should ultimately be the API's user-facing health—not only a healthy scheduler metric.
Logs should survive the Pod when production diagnosis requires history
Kubernetes makes container stdout and stderr available through kubectl logs.
That is useful for immediate diagnosis.
It is not the same as durable cluster-level log storage.
Pods are replaceable. Containers restart. Nodes can disappear.
If production diagnosis requires historical logs, those logs need a lifecycle independent from the Pod or node that produced them.
A useful logging policy answers:
- Which logs are required for incident response?
- How long are they retained?
- Can logs from terminated Pods still be found?
- Are messages structured?
- Is there a request or correlation ID?
- Can the team filter by service, namespace, environment, or severity?
- Are secrets, tokens, credentials, and sensitive payloads excluded?
Structured logging usually matters more than raw log volume.
Prefer predictable fields such as:
timestamp service environment level request_id error_code message
over large volumes of unstructured debug text.
Kubernetes Events are diagnostic evidence, not historical monitoring
Events are especially useful when Kubernetes itself cannot reconcile an object as expected.
Examples:
- FailedScheduling;
- image pull errors;
- volume mount failures;
- node pressure;
- eviction;
- failed probe activity;
- workload lifecycle changes.
Events are useful for questions such as:
Why is this Pod Pending?
Why could this image not start?
Why did this volume fail to attach?
Why was this Pod evicted?
They are less useful for:
What was API latency last Tuesday?
How many checkouts failed yesterday?
Which dependency caused one request to take 900 ms?
Kubernetes Events have limited retention and should be treated as best-effort diagnostic data.
Capture and correlate them when troubleshooting, but do not use Events as a long-term monitoring database.
Traces become valuable when requests cross several services
Distributed traces connect multiple operations into one request path.
Example:
client ↓ 15 ms API ↓ 35 ms internal service ↓ 620 ms database ↓ response
Metrics may tell you p95 latency increased.
Logs may show timeouts.
A trace can show where one instrumented request actually spent its time.
Tracing becomes particularly valuable when:
- requests cross multiple services;
- latency can occur in several dependencies;
- failures are intermittent;
- the system is asynchronous or distributed;
- metrics identify a problem but not the responsible hop.
Kubernetes 1.37 system components can export spans over OpenTelemetry Protocol (OTLP). Application tracing still requires the application and dependencies to emit and propagate trace context.
Do not make tracing the first observability project for a simple application that still lacks basic latency and error-rate metrics.
Start with the simplest signal that closes the current diagnostic gap.
Build a minimum production monitoring baseline
A small Kubernetes team does not need every possible signal on day one.
A useful baseline is:
| Area | Minimum useful signal |
|---|---|
| User impact | availability, error rate, latency |
| Workload health | available replicas, restart loops, failed Jobs |
| Scheduling | persistent Pending Pods and reason |
| Worker health | CPU, memory, node Ready state, pressure |
| Deployments | rollout success/failure |
| Autoscaling | replica changes, Pending Pods, worker scale events |
| Stateful workloads | storage errors and recovery status |
| Application diagnosis | structured logs |
| Kubernetes diagnosis | Events |
| Distributed systems | traces where needed |
The exact list should follow the application.
A background worker may care more about queue age than HTTP latency.
A SaaS API may care more about request success rate and p95/p99 latency.
Monitoring should represent the service the user actually receives.
Alerts should represent action
Do not convert every metric into an alert.
Prioritize alerts by operational consequence.
1. User-visible failures
Examples:
- sustained availability loss;
- error-rate threshold;
- latency beyond the service objective;
- failed critical business operation.
2. Workload failures that threaten service
Examples:
- unavailable replicas;
- repeated CrashLoopBackOff;
- persistent Pending Pods;
- failed critical Job.
3. Capacity conditions with a clear response
Examples:
- sustained memory pressure;
- worker exhaustion;
- autoscaling maximum reached;
- storage nearing a critical limit.
4. Platform conditions that will become incidents
Examples:
- certificate issue;
- repeated node failures;
- storage attach failures;
- control-plane or API degradation where visible.
Every alert should answer:
- Who owns it?
- Why does it matter?
- What is the first diagnostic step?
- What condition closes it?
- How is repeated noise suppressed?
If no action follows from an alert, it probably belongs in a dashboard rather than a notification channel.
Use one troubleshooting path instead of opening every dashboard
Choose the first signal from the symptom.
| Symptom | Start with | Then inspect |
|---|---|---|
| User errors increase | application metrics | logs, traces |
| Latency increases | latency/error metrics | trace path, dependency metrics |
| Pod is Pending | Kubernetes Events | scheduler reason, requests, capacity |
| Pod restarts | restart metric | previous logs, Events, memory/CPU |
| Node pressure | node metrics | requests, evictions, workload distribution |
| Deployment regression | application metric around deploy | rollout state, logs, traces |
| Cluster looks healthy but users fail | service metrics | application logs/traces |
A simple incident flow is:
prove impact → identify scope → inspect Kubernetes state → inspect logs → use traces if distributed → confirm recovery with the original impact signal
The recovery signal should match the failure signal.
If the incident started because checkout errors increased, healthy node CPU does not close the incident.
Checkout errors returning to normal does.
Monitoring and autoscaling share some signals, but have different goals
Autoscaling may use metrics such as CPU or external workload demand.
Monitoring may observe the same metric.
But the intent differs.
Autoscaling asks:
Should capacity change?
Monitoring asks:
Is the system behaving correctly?
Do not assume an autoscaler is a monitoring system.
For HPA and worker scaling, use Kubernetes Autoscaling.
For request sizing that affects both scheduling and HPA utilization, use Kubernetes Requests vs Limits.
Logs and metrics also have cost boundaries
Observability can become a meaningful infrastructure cost.
Cost grows through:
- long retention;
- high-cardinality metrics;
- verbose logs;
- duplicate collection;
- excessive scrape frequency;
- traces with high sampling rates;
- large dashboards querying long time ranges.
The correct answer is not "collect less" by default.
It is to match telemetry cost to operational value.
Examples:
- keep detailed debug logs briefly;
- retain production error logs longer;
- avoid unbounded user IDs as metric labels;
- sample traces where full capture is unnecessary;
- use different retention for production and preview environments.
Use Kubernetes Cost Optimization for the broader infrastructure-cost model.
How Raff Kubernetes monitoring fits this model
Raff Kubernetes currently includes built-in dashboard visibility for:
- live cluster CPU and memory;
- live node CPU and memory;
- Deployments and Pods;
- live Pod logs with tailing;
- cluster activity feed.
Those signals are available without installing an additional observability stack.
Raff also offers one-click Kubernetes applications such as Grafana and related observability tools through its application catalog.
The built-in baseline is useful for questions such as:
- Is a worker under pressure?
- Which workloads are running?
- Is a Pod restarting?
- What is the current container output?
- What cluster activity happened around the failure?
It should not be confused with a complete application-observability system.
Application teams still own:
- service-level metrics;
- business/user-path indicators;
- alert definitions;
- durable log retention;
- application tracing;
- telemetry retention policy.
Use the Raff Kubernetes product page for the current platform capabilities.
Production monitoring checklist
Before calling a Kubernetes workload observable, verify:
Service
- user-impact metrics exist;
- latency/error/availability indicators match the workload;
- alerts have owners and recovery conditions.
Kubernetes
- replica health is visible;
- restart loops are visible;
- Pending Pods can be investigated;
- deployment failures are visible;
- node health and pressure are visible.
Logs
- important application logs are structured;
- sensitive values are excluded;
- retention matches operational need;
- terminated-Pod history is available when required.
Events
- engineers know when to inspect Kubernetes Events;
- Events are not treated as durable historical telemetry.
Traces
- tracing is added where distributed request paths justify it;
- context propagates across the services being investigated;
- trace storage/sampling is bounded.
Operations
- dashboards support diagnosis rather than vanity metrics;
- alerts are actionable;
- incident recovery is confirmed with the original impact signal.
Final recommendation
Start Kubernetes monitoring from operational questions, not tooling.
Track user-facing service health first. Add Kubernetes workload and worker signals to explain capacity and reconciliation problems. Keep structured logs long enough to investigate important failures. Use Events for short-lived Kubernetes diagnostic evidence. Add distributed tracing when request paths become complex enough that metrics and logs cannot explain where time is spent.
Metrics Server and the stable metrics.k8s.io/v1 API provide useful current CPU and memory data, but they are not a complete historical observability system.
On Raff Kubernetes, use the built-in cluster, node, workload, log, and activity visibility as the infrastructure baseline, then add application telemetry according to the incidents your service actually needs to detect and diagnose.
Continue with Kubernetes Cluster Management, Kubernetes Autoscaling, and Kubernetes Cost Optimization for the operational systems around monitoring.