Shared vs dedicated vCPU describes how a cloud VM receives CPU time. Shared vCPU uses host compute scheduled across multiple VMs; dedicated vCPU reserves CPU capacity for one VM so sustained compute behavior is more predictable.
The practical choice is straightforward: use shared vCPU when the workload is bursty, variable, or cost-sensitive and CPU consistency is not the bottleneck. Consider dedicated vCPU when sustained CPU pressure repeatedly affects latency, build time, queue processing, or another measurable workload outcome.
Raff Technologies maps this decision to two VM families: General Purpose uses shared vCPU, while CPU-Optimized uses dedicated compute. This guide focuses on the CPU allocation model. For the broader choice between balanced and CPU-focused machines, start with Cloud VM Machine Classes Explained. If the question is simply whether you need 2 vCPU or 4 vCPU, use 2 vCPU vs 4 vCPU Cloud VMs.
Shared vs dedicated vCPU at a glance
Shared and dedicated vCPU can both run production workloads. The difference is what you are paying to control.
| Factor | Shared vCPU | Dedicated vCPU |
|---|---|---|
| CPU allocation | Host CPU time is scheduled across multiple VMs | CPU capacity is reserved for the VM |
| Main advantage | Lower cost for variable workloads | More predictable sustained CPU behavior |
| Main trade-off | More potential performance variation | Higher cost for reserved compute |
| Strong starting fit | Websites, dev, staging, small apps, variable APIs | CI runners, busy workers, compute-heavy services, latency-sensitive apps |
| Production use | Valid when CPU consistency is not the constraint | Useful when consistency has measurable value |
| Raff family | General Purpose | CPU-Optimized |
A useful mental model is: shared vCPU buys cost efficiency; dedicated vCPU buys CPU predictability.
Neither class fixes insufficient RAM, slow storage, inefficient queries, external API latency, or application code that cannot use additional CPU.
The decision starts with sustained CPU behavior
The decision should connect CPU pressure to a workload outcome.
| Evidence | Better first decision | Why |
|---|---|---|
| CPU is usually low with short bursts | Shared vCPU | Reserved compute would sit unused much of the time |
| Brief deployment or traffic spikes | Stay shared and observe | Short spikes alone do not prove a consistency problem |
| Sustained CPU around 70–80% with rising latency or queue time | Investigate CPU class and size | Compute may be controlling the workload |
| CPU pressure repeats during equivalent builds or jobs | Dedicated vCPU becomes a candidate | Predictable compute can matter to completion time |
| p95/p99 latency rises with sustained CPU pressure | Dedicated vCPU becomes a candidate | Tail latency may be sensitive to CPU scheduling |
| CPU is moderate but RAM is exhausted or swap is active | Add memory or change memory class | CPU allocation is not the primary constraint |
| CPU is low but database queries are slow | Diagnose queries, locks, storage, or connections | Dedicated CPU would not address the main wait |
| Development or staging workload is intermittent | Shared vCPU | Cost efficiency usually matters more than strict consistency |
| Worker or CI workload runs CPU-bound for long periods | Dedicated vCPU | Sustained compute makes reservation more useful |
The 70–80% range is an investigation trigger, not a universal threshold. A workload can operate correctly at high utilization, while another can miss latency goals at a lower average. The important relationship is CPU pressure plus impact.
In Raff infrastructure work, the useful evidence is repeatability. If CPU pressure and user-visible latency or job duration rise together under comparable load, changing CPU class is defensible. If memory, storage, queries, or a dependency controls the outcome, dedicated CPU usually does not fix the problem.

Shared vCPU fits bursty and cost-sensitive workloads
Shared vCPU is usually the sensible default when compute demand is variable and occasional CPU variation does not create a meaningful business or operational problem.
Good fits include:
- websites and content applications with uneven request patterns;
- development and staging environments;
- internal tools with intermittent usage;
- early SaaS applications with modest traffic;
- small APIs that spend substantial time waiting on databases or external services;
- test servers and temporary environments;
- workloads where memory or storage matters more than sustained compute.
Shared CPU is also valid for production when the application is healthy on it. “Production” describes responsibility, not CPU behavior. A low-traffic customer application with comfortable latency and resource headroom does not become CPU-bound merely because users pay for it.
Stay on shared vCPU when the workload meets its latency and throughput targets and CPU pressure is not repeatably controlling the outcome. If more headroom is needed, a larger shared VM may be a simpler and more economical change than moving immediately to dedicated compute.
Use Raff General Purpose VM Plans to compare current shared-vCPU sizes.
Dedicated vCPU fits sustained or latency-sensitive compute
Dedicated vCPU becomes more useful as compute demand becomes sustained, repeatable, and expensive to vary.
Common candidates include:
- CI/CD runners with CPU-bound builds or test suites;
- queue workers processing jobs continuously;
- encoding, compression, compilation, or transformation workloads;
- CPU-heavy application services;
- customer-facing APIs where tail latency correlates with CPU pressure;
- databases whose measured workload is genuinely CPU-bound;
- scheduled processing that must finish inside a predictable window.
The strongest case is a workload where execution time has operational value. If a build blocks every deployment, a worker queue has a deadline, or request latency directly affects users, greater CPU predictability may justify the additional cost.
Dedicated CPU is not automatically the right database class. A database can be limited by memory, query plans, indexes, locks, connections, or storage latency. Confirm CPU is controlling the workload before changing allocation model.
Likewise, a CPU-Optimized VM is not a substitute for application tuning. Reserved compute can make CPU-bound execution more predictable, but it cannot remove work the application should not be doing.

More vCPU and dedicated vCPU solve different problems
A common mistake is treating “more vCPU” and “dedicated vCPU” as the same upgrade.
More vCPU increases available compute capacity. Dedicated vCPU changes how predictably CPU capacity is allocated. A workload may need one, the other, both, or neither.
| Situation | Better first move |
|---|---|
| Shared VM is consistently CPU-saturated | Test a larger shared size or dedicated class |
| CPU utilization is moderate but execution time varies under comparable load | Test dedicated CPU |
| RAM is exhausted or swap is active | Add memory |
| Database is waiting on I/O | Investigate storage and queries |
| App is waiting on external services | Fix dependency behavior before CPU class |
| One service dominates a multi-service VM | Consider separating services |
If the problem is simply insufficient core count, 2 vCPU vs 4 vCPU Cloud VMs provides the narrower sizing framework.
Measure the workload before changing VM class
A dedicated CPU decision should be supported by a before-and-after test rather than a plan label.
Track metrics that connect compute pressure to outcomes:
- CPU utilization over time;
- load or run-queue pressure;
- CPU steal time when the guest exposes it;
- request p50, p95, and p99 latency;
- queue depth and oldest-job age;
- worker throughput and job duration;
- build and test duration;
- database query time;
- memory and swap activity;
- disk I/O wait;
- error and timeout rates.
Compare equivalent workload windows. A production traffic peak should not be compared with an idle overnight period, and a clean build should not be compared with a cached build.
| Test | Shared class | Dedicated class | Decision |
|---|---|---|---|
| Same request mix | Record p95/p99 and CPU pressure | Repeat with equivalent load | Keep dedicated only if the workload outcome improves meaningfully |
| Same CI build | Record duration and variance across runs | Repeat the same pipeline | Look for lower or more predictable duration |
| Same worker batch | Record completion time and queue age | Repeat with the same job set | Confirm throughput improves without moving the bottleneck |
Do not evaluate only average CPU percentage. A class change is successful when latency, completion time, throughput, or operational predictability improves enough to justify the cost.
For a broader diagnostic sequence, use Cloud Server Performance Bottlenecks.