Triage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
During an incident the first job is to stop the harm, not to explain it. Responders lose the most time chasing a root cause while users are still affected, or acting on a guess stated as a fact. Good triage separates what is observed from what is suspected, picks the lowest-risk mitigation that could work, and names the one check that would most change the picture.
Triage this alert: Only if [RECENT_CHANGES] is given: Recent changes: Only if [SIGNALS] is given: Other signals:
- Impact: who is affected (all users, a region, a tenant, an endpoint, internal only), since when, and whether it is getting worse. Say which parts are observed and which are inferred.
- Severity: SEV1 (major user-facing outage or data at risk), SEV2 (significant degradation or a key feature down), SEV3 (minor or partial impact with a workaround), SEV4 (no user impact yet). Give the reason in one line.
- Mitigations: list the options that could stop the harm without knowing the cause, such as rolling back the most recent deploy, turning off a feature flag, failing over, scaling out, shedding or rate-limiting load, or pausing a job. Rank them by how likely they are to help and how risky and reversible they are. A change that lines up in time with the start of the alert goes first.
- Hypotheses: up to four likely causes. For each, the evidence for it, the evidence against it, and the single fastest check that would confirm or rule it out.
- If you have read-only tools (log queries, metrics,
kubectl getordescribe, the repo), run the checks yourself, quote the result, and update the ranking. Ask before anything that changes state. - Escalation: who else to involve now and why (owners of a dependency, the database on-call, communications).
- Only run read-only commands. Never restart, scale, roll back, delete or change configuration yourself; propose it and let the responder run it.
- Never state a root cause as fact. Use "likely", "ruled out" or "confirmed by <evidence>".
- Use UTC timestamps and quote numbers exactly as they appear in the signals.
- Keep it short enough to read in one minute: no background, no generic advice, no restating the alert.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Severity
SEVn: one-line reason.
Impact
Who, since when (UTC) and the trend. Mark each point observed or inferred.
Mitigate now
Numbered, best first. Each: the action, why it might help, its risk, and how to undo it.
Hypotheses
| # | Hypothesis | For | Against | Fastest check |
Next checks
The two or three checks to run next, as exact commands or queries when you know them, with what each result would mean.
Escalate
Who to page or inform, or "Not yet" with the condition that would change it.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Intermediate
- made for
- Site reliability engineer, DevOps / platform engineer, Backend engineer, Software engineer
- needs
- repo-read, shell
- risk
- runs-commands
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install triage-production-alert --target claude-codenpx skills add hermes-hq/hodios-dist --skill triage-production-alert -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsWrite an incident status update
Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.
write-incident-updateWrite a blameless postmortem
Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.
write-postmortemWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observabilityPlan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
plan-game-dayBuild an incident timeline
Builds a timestamped incident timeline from chat logs, alerts and deploy records, marking detection, escalation, mitigation and the gaps between them. Use when preparing a postmortem.
build-incident-timeline