Idle infrastructure is cloud capacity that remains allocated or billable after it stops supporting active production, development, testing, sales, or customer work. For a small team, the cost problem is rarely one large mistake. It is usually several reasonable temporary resources that never receive an end date.
Raff supports 3,000+ customers and 15,000+ VMs. Across that scale, the useful cost-control principle is still simple: current VM prices and plan shapes should be verified on the live pricing page, while dev, staging, preview, test, demo, storage, and recovery resources need explicit ownership and lifecycle rules.
The practical goal is not to delete everything that looks quiet. It is to decide whether each resource should be kept, resized, scheduled, shut down, archived, or deleted based on current business value and recovery risk.
Idle infrastructure has no current job
A resource is idle when its original purpose has ended and nobody can explain why it still needs to remain active or billable.
That is different from an inactive resource. A staging VM may be unused today but intentionally reserved for a release tomorrow. A disaster-recovery resource may spend most of its time quiet but still have a defined role. Idle infrastructure has no active owner, no near-term purpose, or no documented reason to keep paying for it.
| Resource | Legitimate reason to exist | Idle signal |
|---|---|---|
| Dev VM | Active engineering work | Project ended or owner no longer uses it |
| Staging | Release validation or QA | No current release activity and oversized capacity |
| Preview | Open branch or pull request | PR merged or closed |
| Test VM | Bug, load, migration, or integration test | Test cycle completed |
| Demo VM | Scheduled customer or sales use | Demo finished with no follow-up need |
| Volume | Persistent data still required | Unattached with no owner or recovery purpose |
| Snapshot | Active rollback window | Change is stable and retention has no purpose |
| Backup | Recovery requirement | Retention exceeds the defined recovery need |
The invoice is the symptom. Missing lifecycle ownership is the root cause.
Use a keep, resize, schedule, shut down, archive, or delete decision
A useful idle-infrastructure review should end with a specific action.
| Current state | Recommended action | Example |
|---|---|---|
| Active and correctly sized | Keep | Production VM serving customers |
| Still useful but oversized | Resize | Staging copied from production months ago |
| Needed only at known times | Schedule | QA server used during release windows |
| Might be needed soon | Shut down | Developer VM between short projects |
| Server is unnecessary but data may matter | Archive | Old demo database export |
| No owner, purpose, or valuable data | Delete | Preview VM for a merged pull request |
When uncertain, shutting a non-production VM down first is often safer than immediately deleting it. Give the owner a defined review window, then archive or delete according to the value of the data.
Dev environments should default to temporary
Development infrastructure exists to help engineers move quickly. It should not become permanent simply because it was easy to create.
A practical dev-resource policy answers four questions:
- Who owns this resource?
- Which project or task does it support?
- When was it last actively used?
- When will it be reviewed again?
A personal dev VM may be worth keeping when it supports daily work. A temporary build or debugging VM should normally have an expiration date. Shared dev servers need even clearer ownership because they can become dumping grounds for old services and data.
A useful operating rule is: if a dev environment cannot be explained in one sentence, it should not remain online indefinitely.
Staging should match release risk, not production cost
Staging is valuable because it reduces release risk, but it does not automatically need the same capacity or uptime profile as production.
Keep staging always on when teams deploy frequently, QA uses it continuously, integrations depend on it, or recreating the environment would slow delivery. Consider resizing or scheduling it when release activity is infrequent and the environment can be restored predictably.
| Staging pattern | Good fit | Cost-control action |
|---|---|---|
| Always-on | Frequent releases and shared QA | Review size monthly |
| Scheduled | Known test windows | Stop outside active windows |
| Smaller than production | Functional testing without production load | Right-size to staging demand |
| Temporary clone | Migration or major release | Delete after validation |
For broader environment design, use Dev, Staging, and Production Cloud Environments.
Preview environments need automatic endings
Preview environments are intended to be short-lived. Their lifecycle should follow the branch, pull request, feature review, or customer preview that created them.
A simple policy is:
- create when the review begins;
- refresh when the branch changes;
- delete when the PR is merged or closed;
- notify the owner if the preview remains inactive beyond the normal review window.
The ideal control is automated cleanup. If automation is not available, a weekly review is enough for a small team.
A preview environment should usually have a shorter life than the branch that created it. If the branch is gone and the infrastructure remains, the resource is a strong idle-cost candidate.
See Preview Environments vs Staging for the architectural difference between the two patterns.
Test and migration resources need an expiry date
Testing often requires temporarily realistic infrastructure. Load tests may need more compute for a few hours. Migration rehearsals may require a database copy for several days. Bug reproduction can need a dedicated environment until the issue is closed.
Those are valid costs. The problem appears when the test ends and the infrastructure silently becomes part of the permanent monthly baseline.
| Test resource | Useful while | End condition |
|---|---|---|
| Bug reproduction VM | Issue is active | Ticket closed or workaround accepted |
| Load-test server | Performance test scheduled | Test window finished |
| Migration stack | Cutover being validated | Migration completed |
| Integration test VM | External integration under review | Integration accepted or abandoned |
| Temporary database copy | Data migration or QA requires it | Validation completed |
Test environments also create security risk when old credentials, copied customer data, outdated packages, or public admin interfaces survive longer than intended. Cleanup is therefore a cost and security control at the same time.
Demo infrastructure should map to a business reason
Demo environments are often created quickly for prospects, partners, onboarding, or proofs of concept. That urgency makes them easy to forget.
Every demo resource should map to a customer, opportunity, owner, or date. After the meeting or evaluation window, decide whether to keep, archive, or delete it.
A demo VM tied to an active commercial opportunity can be useful infrastructure. A demo VM with no upcoming meeting, no owner, and no customer attached is idle infrastructure.
Storage can remain billable after a VM is gone
Idle infrastructure is not limited to running compute.
Volumes, snapshots, backups, copied databases, logs, and object data can remain after the server that created them has been removed. These resources need the same ownership discipline as VMs.
Use the live Raff pricing and Data Protection pages for current volume, snapshot, backup, and storage allowances rather than hard-coded values in this lifecycle guide.
| Storage item | Keep when | Review when |
|---|---|---|
| Production backup | Recovery requirement exists | Retention exceeds the recovery policy |
| Pre-change snapshot | Rollback window is still active | Change is stable |
| Dev/test snapshot | Rebuild would be expensive | Project closes |
| Unattached volume | Data may still be required | Owner cannot explain its purpose |
| Database copy | Migration, audit, or test need remains | Data becomes stale or sensitive |
The decision is not “backups are expensive, delete them.” It is “which recovery points still protect something worth recovering?”
The hidden cost is operational complexity
Idle resources do more than increase the bill. They also expand the number of systems the team must understand and secure.
Typical hidden costs include:
- old public endpoints and exposed services;
- stale SSH keys or user access;
- monitoring noise from unused systems;
- backup and snapshot clutter;
- confusion during incidents about which resources matter;
- harder cost forecasting because active and abandoned workloads are mixed together.
This is why lifecycle control becomes more important as the account grows. With five resources, cleanup feels optional. With dozens, nobody wants to delete anything because nobody is sure what is safe to remove.
