# Runbook: service or failure mode

- Owner: Team
- Severity: SEV level
- Last exercised: YYYY-MM-DD

## Trigger

Describe the alert, symptom, and impact threshold.

## Safety checks

- Confirm the environment and tenant scope.
- Preserve request, trace, and job identifiers.
- Do not copy credentials or private content into tickets.

## Diagnose

1. Check liveness and readiness.
2. Inspect privacy-safe metrics and recent deployments.
3. Bound the affected time window and workload.

## Contain

State reversible, authorized containment steps and stop conditions.

## Recover

Document the ordered recovery procedure and rollback.

## Verify

- [ ] Health and readiness are green.
- [ ] A representative user workflow succeeds.
- [ ] Backlog and error rate return to expected bounds.

## Communicate

Name the incident channel, status update owner, and update interval.

## Follow-up

Record evidence, root cause, corrective actions, and the next exercise date.
