Write an incident status update
Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.
During an incident, people judge the team by its updates as much as by the fix. Good updates are early, specific about who is affected, honest about what is not yet known, and regular. Bad ones guess at causes, blame a vendor, promise times the team cannot meet, hide behind jargon or go silent for an hour, and each of those costs trust that is hard to win back. Updates are written under time pressure, so draft immediately instead of asking questions.
Write a update for from these facts: Only if [NEXT_UPDATE] is given: Next update:
- Lead with the impact in the reader's terms: what they cannot do, since when (UTC), and who is affected. Say what still works when the facts show it.
- Say what the team is doing now, matching the phase: investigating (looking into it), identified (cause found, fix under way; describe the cause only in general terms and only if the facts confirm it), monitoring (fix applied, watching, what users may still see), resolved (back to normal, the duration with start and end times, anything users need to do, and a pointer to a follow-up review if one is planned).
- Include a workaround only if the facts contain one.
- End with when the next update will come. If no time was given and the phase is not resolved, use 30 minutes after the current time for investigating and identified, and 60 minutes for monitoring; if the current time is not in the facts either, add
[next update time]for the author to fill in. - If a must-have fact is missing (what is affected, or since when), still write the draft, insert
[CONFIRM: what is needed]at that spot, and list it under Held back. - Tune it to the audience:
- customers: plain language, no internal system names, at most 120 words.
- internal: the affected services, the incident channel or commander if given, the customer impact in numbers if known, what other teams should and should not do, and a suggested line for support to give customers, at most 150 words.
- executives: business impact first (customers, revenue, SLA, regulatory exposure if the facts mention it), the decision or support needed from them if any, at most 100 words.
- Use only the facts given. Never guess a cause, a number of affected users or a resolution time.
- Do not blame a vendor, a team or a person.
- Do not promise a fix time unless the facts contain one the team has committed to.
- Do not apologise more than once, and do not use filler such as "we take this very seriously".
- Times in UTC unless the facts use another timezone. No emoji, no exclamation marks, no marketing language.
Title
One line, for a status page or subject line, stating the affected feature and the phase.
Update
The message, ready to paste.
Short version
Under 280 characters, for an in-app banner or social post.
Held back
Bullets: facts from the input you left out for this audience and why, plus every [CONFIRM] or other placeholder the author must fill before posting. "Nothing" if empty.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Beginner
- made for
- Site reliability engineer, Engineering manager, Customer support, Product manager
- risk
- read-only
- version
- v1.1.0 · experimental
- reviewed
- 2026-10-02
- aliases
- write-status-page-update
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install write-incident-update --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-incident-update -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsTriage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
triage-production-alertWrite a blameless postmortem
Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.
write-postmortemIncident commander
Runs a live incident like an experienced incident commander, assigning roles, keeping a steady comms cadence and driving mitigation before root cause. Use as the coordinating voice during an outage.
incident-commanderInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observabilityPlan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
plan-game-dayBuild an incident timeline
Builds a timestamped incident timeline from chat logs, alerts and deploy records, marking detection, escalation, mitigation and the gaps between them. Use when preparing a postmortem.
build-incident-timeline