hermes

Write an operational runbook

Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.

context

A runbook is read by a tired engineer who may never have touched this system, often in the middle of the night. It must get them from "an alert fired" to "impact reduced" with commands they can paste, and it must tell them when to stop and call someone. Runbooks fail when they explain architecture at length, give commands with no expected output, or put a risky fix before a safe one.

task

Write a runbook for: Only if [SYSTEM_CONTEXT] is given: System context:

  1. Decide which kind this is. For an alert, write the alert flow below. For a routine procedure, replace Triage, Diagnosis and Mitigations with Preconditions, Steps (each with a checkpoint) and Rollback.
  2. Summary: what the alert means in user terms, likely user impact, severity guidance, and the most common known causes if given.
  3. Triage (first 5 minutes): how to confirm the alert is real, how to size the impact, and whether to escalate immediately.
  4. Diagnosis: read-only checks in order of likelihood. Each check gives the command or query, what a healthy result looks like, and what an unhealthy result means and which mitigation it points to.
  5. Mitigations: ordered from safest and most reversible to riskiest. Each states when to use it, the exact steps, the risk, and how to undo it.
  6. Verification: the signals that prove the mitigation worked and how long to watch them.
  7. Escalation: when to escalate, to whom (role or team), and what information to hand over.
constraints
  • Never invent hostnames, dashboard links, metric names, namespaces or team names. Use placeholders in angle brackets such as <service-namespace> and list every one under "Fill before publishing".
  • Put every command in a fenced block. Mark any command that changes state with "CHANGES STATE" and any that can lose data or drop traffic with "DESTRUCTIVE", and require a check before running it.
  • Keep it scannable: numbered steps, short sentences, no history lessons.
  • Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
  • Keep the change as small as it can be while still being correct.
output format

Summary

Triage

Diagnosis

Mitigations

Verification

Escalation

Fill before publishing

A checklist of every placeholder and unconfirmed assumption.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Incident and operations
level
Intermediate
made for
Site reliability engineer, DevOps / platform engineer, Backend engineer
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install write-runbook --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill write-runbook -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PersonaIncident and operations

Incident commander

Runs a live incident like an experienced incident commander, assigning roles, keeping a steady comms cadence and driving mitigation before root cause. Use as the coordinating voice during an outage.

incident-commander
PromptIncident and operations

Instrument a service for observability

Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.

instrument-service-observability
PromptIncident and operations

Plan a game day or chaos exercise

Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.

plan-game-day
PromptIncident and operations

Triage a production alert

Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.

triage-production-alert
PromptIncident and operations

Write an incident status update

Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.

write-incident-update
PromptIncident and operations

Write a blameless postmortem

Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.

write-postmortem