hermes

Write a blameless postmortem

Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.

context

A postmortem exists so the same incident does not happen again and the next one is handled faster. That only works when people can describe what they did without fear, so the document explains how the system and its processes allowed a reasonable action to cause harm. "Human error" is where the analysis starts, not where it ends.

task

Write a postmortem from these notes:

  1. Build the timeline first, in UTC, from the notes only. Mark the key moments: start of impact, detection, response start, mitigation, resolution. Compute time to detect, time to mitigate and total duration from them.
  2. Quantify the impact from the notes: users or requests affected, error rates, data lost or delayed, money or SLA effects. Use the notes' numbers only.
  3. Explain the contributing factors as a chain: the trigger, the conditions that let it cause harm, and why detection or mitigation took as long as it did. There is usually more than one factor; list each.
  4. Note what went well, what was hard, and where the team got lucky.
  5. Propose action items, at most seven, each tied to a contributing factor and typed as prevent, detect or mitigate. Each must be specific enough that someone could tell when it is done.
  6. For a public audience, drop internal names, hostnames, tools and people. Keep the impact, the cause in plain words, and the commitments.
constraints
  • Never invent a timestamp, number or event. Write [unknown] and add the gap to Open questions.
  • Blameless language: describe actions, decisions and system conditions, not people's character or competence. Refer to people by role ("the on-call engineer"), never by name.
  • Do not name a single root cause when the notes show several factors.
  • No vague action items such as "be more careful" or "improve monitoring". Name the alert, test, limit or process change.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Summary

Three sentences: what happened, the impact, and how it was resolved.

Impact

Bullets with numbers, duration, and who was affected. Then time to detect, time to mitigate and total duration.

Timeline

| Time (UTC) | Event | Key moments in bold.

Contributing factors

Numbered, starting with the trigger.

What went well

Bullets.

What was hard

Bullets, including where the team got lucky.

Action items

| # | Action | Type (prevent / detect / mitigate) | Factor | Priority | Owner | Leave Owner as TBD.

Open questions

Gaps in the notes that the team should fill in. "None" if empty.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Incident and operations
level
Intermediate
made for
Site reliability engineer, Engineering manager, Software engineer, DevOps / platform engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install write-postmortem --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill write-postmortem -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptIncident and operations

Triage a production alert

Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.

triage-production-alert
PromptIncident and operations

Write an incident status update

Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.

write-incident-update
PromptIncident and operations

Instrument a service for observability

Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.

instrument-service-observability
PromptIncident and operations

Plan a game day or chaos exercise

Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.

plan-game-day
PromptIncident and operations

Build an incident timeline

Builds a timestamped incident timeline from chat logs, alerts and deploy records, marking detection, escalation, mitigation and the gaps between them. Use when preparing a postmortem.

build-incident-timeline
PromptIncident and operations

Define SLOs and burn-rate alerts

Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.

define-slos