Kubernetes requests and limits control compute resources for Pods and containers. A request tells the scheduler how much CPU or memory to plan for; a limit sets a runtime boundary on how much of that resource a container may consume.
These settings are not API rate limits, ingress rate limits, or request-per-second throttles. In Kubernetes resource management, “requests vs limits” means CPU, memory, and related compute-resource policy.
Raff Technologies supports managed Kubernetes for teams that want to run containerized workloads without operating the control plane themselves. Accurate requests and sensible limits matter because they affect scheduling, node capacity, autoscaling signals, performance, and failure behavior.
Kubernetes requests vs limits: quick answer
| Setting | Scheduler uses it? | Runtime effect | Common failure when wrong |
|---|---|---|---|
| CPU request | Yes | Influences CPU weighting under contention | Over-requesting wastes schedulable capacity; under-requesting packs workloads too tightly |
| Memory request | Yes | Helps represent expected working-set needs | Under-requesting can increase node memory pressure |
| CPU limit | No for placement | CPU usage above the ceiling is throttled | Latency and throughput can degrade during legitimate bursts |
| Memory limit | No for placement | Memory use is constrained; exceeding the boundary can result in OOM termination | Restarts or instability when the limit is below real peak memory |
The simplest rule is:
Requests = scheduling baseline Limits = runtime boundary
Requests answer “what capacity should Kubernetes plan for?” Limits answer “how far may this container consume after it is running?”
Kubernetes resource requests are scheduling inputs
The Kubernetes scheduler places Pods using resource requests rather than waiting to observe their future CPU or memory usage.
For example:
resources: requests: cpu: "250m" memory: "256Mi" limits: cpu: "500m" memory: "512Mi"
This container requests:
250mCPU, or 0.25 CPU;256Mimemory.
It is limited to:
500mCPU;512Mimemory.
If four Pods each request 500m CPU, their combined CPU request is:
4 × 500m = 2000m = 2 CPU
That requested capacity matters before actual runtime usage is known.
Requests that are too high
Inflated requests can:
- make Pods unschedulable earlier;
- leave node capacity unused;
- cause a cluster to appear full even when observed usage is low;
- contribute to unnecessary node scale-out when cluster autoscaling is used.
Requests that are too low
Understated requests can:
- pack too many active workloads onto one node;
- increase CPU or memory contention;
- make capacity planning misleading;
- distort resource-utilization autoscaling signals.
A request should therefore represent a defensible operating baseline, not an arbitrary “safe” number.
For broader worker sizing, see Kubernetes Cluster Sizing.
CPU requests and CPU limits solve different problems
CPU resources are compressible: when a workload cannot get as much CPU as it wants, it can usually continue running more slowly.
A CPU request influences scheduling and relative CPU allocation under contention. A CPU limit defines a ceiling on CPU time the container may consume.
If a container tries to use CPU beyond its configured limit, Kubernetes relies on the container runtime and kernel mechanisms to throttle that workload.
That creates an important production trade-off:
Low CPU limit -> stronger containment -> greater risk of throttling legitimate bursts Higher/no CPU limit -> more burst freedom -> less strict noisy-neighbor containment
For latency-sensitive web applications, a CPU limit that looks generous during average load can still cause problems during short bursts, garbage collection, startup, or traffic spikes.
Do not assume that every CPU limit should equal the request or that every limit should be exactly 2× the request. Measure burst behavior first.
Memory requests and memory limits have different failure behavior
Memory is not handled like CPU.
A memory request influences scheduling. A memory limit constrains how much memory a container is allowed to use.
Memory-limit enforcement is reactive rather than gradual CPU-style throttling. When a process attempts to use memory beyond its allowed boundary, the kernel can terminate a process in the container, producing an out-of-memory failure such as OOMKilled.
Operationally:
CPU exceeds limit -> throttling / slower execution Memory exceeds usable boundary -> process may be terminated
That makes memory sizing especially important for applications with:
- garbage-collected runtimes;
- large startup peaks;
- caches;
- image/video processing;
- batch jobs;
- unpredictable query buffers;
- memory leaks.
A memory request should reflect the workload's normal working set closely enough for responsible placement. A memory limit should reflect how much growth is acceptable before termination is preferable to further node pressure.
Kubernetes CPU and memory units are easy to misread
CPU is measured in CPU units.
1 CPU = 1 CPU core or 1 virtual core 500m = 0.5 CPU 250m = 0.25 CPU 100m = 0.1 CPU
100m means one hundred millicpu, not 100 CPUs.
Memory is expressed as a quantity such as:
256Mi 512Mi 1Gi
Do not reuse CPU-style intuition for memory units. A value such as 400m in a memory field is not 400 MiB.
A small unit typo can turn a reasonable manifest into a workload that fails immediately or is scheduled very differently from what the author intended.
What happens if you set a limit but no request?
This is a common source of confusion.
When a container has a limit but no explicit request for that resource, Kubernetes may use the limit value as the request unless another admission-time defaulting mechanism supplies a value.
For example:
resources: limits: cpu: "1"
can result in an effective CPU request of 1 as well.
That means a team intending only to cap CPU usage can accidentally create a much larger scheduling request than expected.
Namespace policy can also inject values through a LimitRange, so always inspect the effective manifest and runtime configuration, not only the source YAML you remember writing.
Requests influence Horizontal Pod Autoscaler utilization
Requests are also important when Horizontal Pod Autoscaler uses resource utilization targets.
For CPU utilization, a simplified model is:
CPU utilization = current CPU usage / CPU request
Suppose a Pod is using 200m CPU.
With a 100m request:
200m / 100m = 200% utilization
With a 500m request:
200m / 500m = 40% utilization
The workload is doing the same amount of CPU work, but the autoscaling signal is completely different.
Kubernetes also documents that if the relevant resource request is missing for containers in a targeted Pod, resource-utilization calculation can be undefined and the autoscaler may not act on that metric as expected.
This is why HPA tuning should not begin with the HPA target alone. First make resource requests credible.
Requests and limits affect Pod Quality of Service
Kubernetes classifies Pods into QoS classes including:
Guaranteed;Burstable;BestEffort.
At a high level:
- a Pod with no CPU or memory requests/limits can fall into
BestEffort; - a Pod with some resource configuration commonly falls into
Burstable; Guaranteedrequires strict CPU and memory request/limit relationships for the applicable resources and containers.
QoS matters during node resource pressure because eviction behavior differs between classes.
Do not optimize resource settings purely to obtain a QoS label. Start with the workload's real operating requirements, then understand the QoS class that results.
LimitRange and ResourceQuota are not the same thing
These two Kubernetes policies solve different problems.
LimitRange
A LimitRange can apply namespace-level constraints and defaults for individual objects, including:
- default requests;
- default limits;
- minimum values;
- maximum values;
- maximum limit-to-request ratios.
It is useful as a guardrail when workload manifests may omit resource configuration.
ResourceQuota
ResourceQuota governs aggregate namespace consumption and object counts.
The distinction is:
LimitRange -> defaults and boundaries for individual workloads/objects ResourceQuota -> aggregate namespace consumption
Shared clusters may use both.
Neither replaces proper workload sizing. Defaults should catch omissions, not force every service into the same resource shape.
Requests vs limits best practices for production
There is no universal request-to-limit ratio that works for every application.
Start from observed behavior.
| Workload behavior | Request approach | Limit approach |
|---|---|---|
| Steady CPU with short bursts | Request near sustained demand | Leave measured burst headroom if policy allows |
| Bursty latency-sensitive API | Avoid requesting every peak | Test whether CPU limit causes throttling during valid bursts |
| Stable memory working set | Request near normal working set with justified headroom | Set a tested boundary above normal/expected peaks |
| Startup memory spike | Include startup behavior in measurements | Ensure limit survives legitimate startup peak |
| Batch job | Request realistic baseline | CPU limit can protect neighbors if slower completion is acceptable |
| Memory leak | Fix the application rather than masking it with a large request | Limit can provide containment while the root cause is fixed |
| Critical service | Request enough capacity to reflect importance | Choose boundaries from failure policy, not a generic ratio |
Review several operating windows:
- normal traffic;
- peak traffic;
- application startup;
- rolling deployment overlap;
- scheduled jobs;
- failure recovery;
- unusual but legitimate bursts.
For each workload, record:
- typical CPU usage;
- peak CPU and burst duration;
- normal memory working set;
- peak memory;
- startup memory;
- replica count;
- CPU-throttling impact;
- OOM/restart impact.
Common Kubernetes requests and limits mistakes
Mistake 1: copying values from another application
Two services with the same language and framework can have completely different resource behavior.
Mistake 2: assuming request equals average usage
A request is a scheduling commitment. Average usage can be useful evidence, but startup peaks, sustained load, contention, and failure policy also matter.
Mistake 3: setting every CPU limit very close to the request
This can remove useful burst capacity and create unnecessary throttling.
Mistake 4: treating memory limits like CPU limits
Memory cannot simply be slowed down when the boundary is exceeded. OOM failure behavior must be tested.
Mistake 5: using limits but forgetting requests
This can create unexpected effective requests and misleading scheduling behavior.
Mistake 6: confusing resource limits with traffic rate limiting
Kubernetes CPU/memory requests and limits do not control HTTP requests per second, API quotas, or ingress traffic rates. Those are separate networking/application concerns.
Mistake 7: tuning HPA before tuning requests
Utilization-based HPA depends on requests. Bad requests produce bad percentages.
How requests affect worker capacity and cost
Resource requests directly influence whether new Pods fit on existing workers.
The chain is straightforward:
Pod requests ↓ Scheduler evaluates node allocatable resources ↓ Pod fits or remains pending ↓ Cluster capacity may need to change
Inflated requests can make a cluster appear capacity-constrained earlier than necessary. Requests that are too low can produce tighter packing and contention after scheduling.
This is why requests are also a cost-management input, not just a manifest detail.
For Raff Kubernetes, use the live Kubernetes product page and pricing page for current worker sizes, autoscaling capabilities, control-plane availability, and pricing rather than relying on static article numbers.
For the broader cost model, continue with Kubernetes Cost Optimization.
How to review requests and limits over time
Resource policy should change when the workload changes.
Revisit settings after:
- major releases;
- traffic changes;
- runtime upgrades;
- new background jobs;
- database-client changes;
- replica-strategy changes;
- repeated CPU throttling;
- repeated OOM kills;
- node-pressure incidents;
- HPA behavior that does not match expectations.
A recurring review should ask:
- Are requests close to sustained useful demand?
- Are Pods pending because capacity is truly insufficient or because requests are inflated?
- Is CPU throttling correlated with latency or job duration?
- Are memory limits causing legitimate workloads to restart?
- Are memory requests representative of the normal working set?
- Are HPA targets calculated against credible requests?
- Are namespace defaults still appropriate?
- Does this workload need a different node pool or isolation boundary?
Conclusion
Kubernetes requests vs limits is fundamentally a scheduling-versus-runtime decision.
Requests tell Kubernetes how much resource capacity to plan for when placing workloads. Limits define how much resource consumption is allowed after those workloads are running. CPU and memory behave differently: CPU can be throttled, while memory pressure can result in process termination.
Start with measurements, keep resource policy separate from traffic rate limiting, and review the values as applications evolve. Accurate requests improve scheduling and autoscaling signals; sensible limits help contain runtime behavior without turning normal bursts into avoidable incidents.