Write an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
A runbook is read by a tired engineer who may never have touched this system, often in the middle of the night. It must get them from "an alert fired" to "impact reduced" with commands they can paste, and it must tell them when to stop and call someone. Runbooks fail when they explain architecture at length, give commands with no expected output, or put a risky fix before a safe one.
Write a runbook for: Only if [SYSTEM_CONTEXT] is given: System context:
- Decide which kind this is. For an alert, write the alert flow below. For a routine procedure, replace Triage, Diagnosis and Mitigations with Preconditions, Steps (each with a checkpoint) and Rollback.
- Summary: what the alert means in user terms, likely user impact, severity guidance, and the most common known causes if given.
- Triage (first 5 minutes): how to confirm the alert is real, how to size the impact, and whether to escalate immediately.
- Diagnosis: read-only checks in order of likelihood. Each check gives the command or query, what a healthy result looks like, and what an unhealthy result means and which mitigation it points to.
- Mitigations: ordered from safest and most reversible to riskiest. Each states when to use it, the exact steps, the risk, and how to undo it.
- Verification: the signals that prove the mitigation worked and how long to watch them.
- Escalation: when to escalate, to whom (role or team), and what information to hand over.
- Never invent hostnames, dashboard links, metric names, namespaces or team names. Use placeholders in angle brackets such as
<service-namespace>and list every one under "Fill before publishing". - Put every command in a fenced block. Mark any command that changes state with "CHANGES STATE" and any that can lose data or drop traffic with "DESTRUCTIVE", and require a check before running it.
- Keep it scannable: numbered steps, short sentences, no history lessons.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
Summary
Triage
Diagnosis
Mitigations
Verification
Escalation
Fill before publishing
A checklist of every placeholder and unconfirmed assumption.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Intermediate
- made for
- Site reliability engineer, DevOps / platform engineer, Backend engineer
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install write-runbook --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-runbook -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsIncident commander
Runs a live incident like an experienced incident commander, assigning roles, keeping a steady comms cadence and driving mitigation before root cause. Use as the coordinating voice during an outage.
incident-commanderInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observabilityPlan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
plan-game-dayTriage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
triage-production-alertWrite an incident status update
Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.
write-incident-updateWrite a blameless postmortem
Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.
write-postmortem