A runbook is a repeatable operating guide for a known production event: an incident, failed deployment, access change, patch window, restore, rebuild, or another task where the wrong action can create downtime or data loss.
For a small team, a good runbook should answer seven things quickly: what triggered it, who owns the response, what to check first, which actions are safe, which actions need approval, how to recover or roll back, and what proves the system is healthy again.
Raff Technologies workloads can run across Linux VMs, Windows VMs, private networks, and protected storage. That flexibility is useful only when the team knows how to operate the environment repeatedly. A runbook turns infrastructure knowledge into a process another engineer can execute under pressure.
Runbook quick answer
Use a runbook when an operational event is repeatable, time-sensitive, and risky enough that improvisation is dangerous.
| Event | Recommended runbook |
|---|---|
| Production outage | Incident response runbook |
| Failed release | Deployment rollback runbook |
| Leaked credential or admin change | Access and credential-rotation runbook |
| OS/package maintenance | Patch and rollback runbook |
| Data loss or corruption | Restore/recovery runbook |
| VM becomes untrusted or unrecoverable | Rebuild runbook |
| RDP/IIS/Windows service issue | Windows operations runbook |
| SSH/systemd/Linux service issue | Linux operations runbook |
The purpose is not to document every command. It is to make the decision path clear before the team is under pressure.
What is a runbook?
A runbook is an operational document that describes how to handle a known event safely and consistently.
It differs from general documentation:
| Document | Main purpose |
|---|---|
| Architecture documentation | Explains how the system is designed |
| Setup guide | Explains how to install or provision something |
| Runbook | Explains what to do when a specific operational event occurs |
| Playbook | Coordinates a broader scenario across people, systems, or teams |
| Checklist | Confirms required steps or conditions |
| Post-incident review | Explains what happened and what should change |
A runbook is most useful when it contains decision gates, not only steps.
“Restart the server” is a weak instruction. “Restart only if liveness has failed, no data operation is in progress, and restart is expected to improve the failure mode” is a real operational decision.
Runbook template for small teams
Use this template as the default structure for a cloud operations runbook.
| Section | What to document |
|---|---|
| Runbook name | Clear event or task |
| Applies to | Service, VM, environment, or workload |
| Owner | Person or role responsible for accuracy |
| Trigger | Conditions that make the runbook relevant |
| Impact/severity | Expected customer or business impact |
| First checks | Evidence needed before taking action |
| Safe actions | Low-risk steps allowed immediately |
| Approval gates | Destructive or high-risk steps requiring explicit approval |
| Rollback/recovery | How to reverse or recover from the action |
| Verification | Evidence that the system is healthy again |
| Escalation | Who to involve when the runbook no longer fits |
| Communication | Who needs updates and when |
| Post-action notes | What to record after execution |
| Last reviewed | Date and owner of the last review |
A useful runbook is often short. The responder should be able to find the trigger, first action, and stop condition without reading an essay.
The most important field is the trigger
A runbook should begin with a specific condition that tells the team when to use it.
Weak triggers create ambiguity:
| Weak trigger | Better trigger |
|---|---|
| “Server is broken” | “Production API returns sustained 5xx errors and customer requests are failing” |
| “Database issue” | “Connection acquisition or query latency is blocking production requests” |
| “Deployment failed” | “New release causes elevated errors, failed readiness checks, or a rollback decision” |
| “Access problem” | “Admin credential, SSH key, API token, or account must be granted, removed, or rotated” |
| “Backup issue” | “Restore is required or a recovery point cannot be validated” |
A metric alone is not always a trigger. High CPU can be normal if the workload is productive. The runbook should be tied to an operational decision or customer impact.
Incident response runbook
An incident response runbook gives the team a predictable first response when production is degraded or unavailable.
NIST SP 800-61 Rev. 3 treats incident response as part of a broader risk-management lifecycle that includes preparation, detection, response, recovery, and improvement. Google SRE guidance similarly emphasizes clear ownership, communication, and coordination during incidents.
A small-team incident runbook should include:
| Section | Decision |
|---|---|
| Declare | Does this meet the team's incident threshold? |
| Severity | What is the customer/business impact? |
| Owner | Who coordinates the response? |
| Triage | What is affected, how broadly, and since when? |
| Evidence | Which logs, metrics, timestamps, and changes must be preserved? |
| Containment | What can reduce impact without destroying useful evidence? |
| Recovery | Restart, rollback, restore, rebuild, or fail over? |
| Communication | Who needs an update and when? |
| Exit criteria | What proves the incident is resolved? |
| Review | What should change afterward? |
One person may hold multiple roles in a small team. The important part is that ownership is explicit.
See Incident Response Plan for Small Teams for the broader response framework.
Deployment runbook
A deployment runbook defines how a release moves from “ready to deploy” to “verified in production,” including when to stop or roll back.
Minimum sections:
- deployment owner;
- change summary;
- risk level;
- active incidents or maintenance conflicts;
- database/schema migration risk;
- backup or restore prerequisite where appropriate;
- rollout sequence;
- health/readiness checks;
- rollback or fix-forward path;
- stop condition;
- post-deployment validation.
A practical release decision should look like this:
Prechecks pass -> deploy to limited scope -> readiness succeeds -> critical user path succeeds -> error/latency signals remain acceptable -> continue rollout
If those checks fail, the runbook should already define whether the team rolls back, pauses, or fixes forward.
If rollback is impossible—such as after an irreversible data migration—the runbook must state the alternative recovery path before the change begins.
Access and credential runbook
Access changes deserve a runbook because they combine security risk with the possibility of locking the team out of production.
Common triggers include:
- new engineer onboarding;
- employee or contractor offboarding;
- compromised SSH key;
- leaked API key;
- Windows administrator change;
- emergency production access;
- secret or password rotation.
A safe access runbook should document:
- Who approves the change.
- Which systems are affected.
- Whether access is temporary or permanent.
- Which dependent services use the credential.
- How the old credential is revoked.
- How new access is tested before the old path disappears.
- What audit evidence is retained.
For emergency access, define an expiration or review step. Temporary incident access should not silently become permanent production access.
Recovery and disaster recovery runbooks
A backup is not a runbook. The restore process is the runbook.
A recovery runbook should answer:
- What event requires restore or rebuild?
- What is the acceptable Recovery Point Objective (RPO)?
- What is the Recovery Time Objective (RTO)?
- Which recovery point is trusted?
- Who approves the restore?
- In what order do dependencies recover?
- What proves the recovered data is correct?
- When can customer traffic return?
A basic dependency-aware sequence can be:
Administrative access -> private/network path -> database and persistent data -> application service -> internal dependencies -> workers and scheduled jobs -> public traffic
Do not bring background jobs back before their data dependencies are ready. Retry storms and duplicate side effects can turn recovery into another incident.
See Disaster Recovery Plan for Small Teams and Cloud Server Backup Strategy for deeper recovery planning.
Patch and maintenance runbook
Routine patching can still create downtime.
A patch runbook should define:
| Decision | What to record |
|---|---|
| Urgency | Routine, urgent, emergency |
| Scope | Which systems are affected |
| Window | When maintenance occurs |
| Owner | Who applies and verifies the change |
| Prechecks | Current health, access, recovery path |
| Expected impact | Reboot, service restart, downtime, none |
| Validation | Service status, application health, logs |
| Rollback | Restore, rebuild, package rollback, app rollback |
| Deferral | What compensating control applies if patching waits |
Use Cloud VM Patch Management for the maintenance-window and rollback decision model.
Rebuild runbook: when repair is the wrong choice
Sometimes repairing a server is slower or less trustworthy than rebuilding it.
A rebuild runbook is useful when:
- the operating system is badly damaged;
- a security incident makes host trust uncertain;
- configuration drift is severe;
- storage or filesystem damage makes recovery unreliable;
- replacement from known-good configuration is faster than diagnosis.
The runbook should define:
- What data must be preserved before replacement.
- Which configuration is authoritative.
- How a replacement VM is provisioned.
- How private networking and firewall paths are restored.
- How persistent data is reattached or restored.
- How the application is deployed.
- How customer traffic moves to the replacement.
- When the old server can be destroyed.
The last step matters. Do not delete the original system until required evidence, data, and rollback options have been preserved.
Windows and Linux runbooks need different operational details
The decision structure can be shared, but Windows and Linux operating steps often differ.
| Linux examples | Windows examples |
|---|---|
| SSH access | RDP access |
| systemd services | Windows Services |
| journald/syslog | Event Viewer |
| apt/dnf package changes | Windows Update / role-specific maintenance |
| Nginx/Apache | IIS |
| shell scripts | PowerShell |
For Windows workloads, document how the team validates RDP, IIS/application services, Event Viewer findings, patching, administrator access, and any workload-specific licensing requirements.
For Linux workloads, document service names, log locations, firewall assumptions, deployment paths, and which directories or files must never be removed casually.
The goal is not two separate operations cultures. It is to remove OS-specific tribal knowledge from the incident path.
Runbook automation: when should you automate?
Runbook automation is useful after the team understands the procedure well enough to define safe inputs, failure conditions, approvals, and rollback behavior.
Automate when:
- the task repeats frequently;
- the decision rules are measurable;
- the inputs are predictable;
- the action is idempotent or safely repeatable;
- rollback is understood;
- human approval can be placed before destructive steps.
Keep a human decision when:
- data loss is possible;
- the failure mode is ambiguous;
- the action is difficult to reverse;
- business context determines the correct choice;
- the system changes faster than the automation.
A useful progression is:
Document -> execute manually -> test repeatedly -> add checks and guardrails -> automate low-risk steps -> preserve approval for destructive actions
Do not automate a vague runbook. Automation only makes a bad procedure execute faster.
See Automation & Infrastructure-as-Code on Raff for infrastructure repeatability beyond incident operations.
Verification and exit criteria make a runbook trustworthy
A runbook should never end with “command completed.” It should end with evidence.
Useful verification signals include:
- health/readiness endpoints;
- critical user journeys;
- application error rate;
- latency;
- queue depth or oldest-job age;
- database connectivity and query success;
- logs showing expected state;
- CPU, memory, disk, and network saturation;
- backup/replication status where relevant.
Define the exit criteria before the incident if possible.
For example:
“Recovery is complete when login succeeds, critical writes complete, error rate is back within normal range, background jobs are moving, and no critical alerts remain.”
That is stronger than “the VM is online.”
Review runbooks after real use
A stale runbook can be dangerous because responders assume it is correct.
Review a runbook when:
- it was used during an incident;
- a deployment failed;
- a restore test exposed missing steps;
- access or secret management changed;
- the architecture changed;
- a service moved to another host or dependency;
- the team discovered an unclear instruction.
Instead of changing a date simply to look fresh, record a review when someone has actually verified the procedure.
Useful ownership fields are:
- owner;
- last reviewed date;
- systems covered;
- known limitations;
- change history.
Minimum runbook set for a small team
A small team does not need dozens of runbooks immediately.
Start with the operational failures most likely to create serious business impact:
- Production incident response — who owns the outage and how triage begins.
- Deployment rollback — how to stop a bad release from spreading.
- Access onboarding/offboarding — how privileged access is granted and removed.
- Backup restore — how a recovery point becomes a working service.
- Patch maintenance — how updates are applied and reversed.
- Server rebuild — how an untrusted or damaged VM is replaced.
- OS-specific operations — the Windows or Linux actions your workload actually uses.
Add more only when repeated incidents or operational risk justify them.
Applying runbooks on Raff
Raff can provide the infrastructure building blocks; the team still owns the operating decisions in its runbooks.
A practical model is:
| Operational need | Raff building block |
|---|---|
| Application or service host | Raff Cloud VM |
| Linux workload | Linux VM |
| Windows workload | Windows VM |
| Private service communication | VPC |
| VM backup and recovery planning | Data Protection |
For a production application, document the exact access path, service names, deployment process, health checks, backup/recovery path, and escalation contact before the system becomes business-critical.