Runbook: service or failure mode
- Owner: Team
- Severity: SEV level
- Last exercised: YYYY-MM-DD
Trigger
Describe the alert, symptom, and impact threshold.
Safety checks
- Confirm the environment and tenant scope.
- Preserve request, trace, and job identifiers.
- Do not copy credentials or private content into tickets.
Diagnose
- Check liveness and readiness.
- Inspect privacy-safe metrics and recent deployments.
- Bound the affected time window and workload.
Contain
State reversible, authorized containment steps and stop conditions.
Recover
Document the ordered recovery procedure and rollback.
Verify
- Health and readiness are green.
- A representative user workflow succeeds.
- Backlog and error rate return to expected bounds.
Communicate
Name the incident channel, status update owner, and update interval.
Follow-up
Record evidence, root cause, corrective actions, and the next exercise date.