hermes

Design actionable alerting rules

Designs actionable alerts from SLOs and user-facing symptoms, with thresholds, routing, runbook links, and a list of noisy alerts to delete. Use when pages are noisy or real outages go unnoticed.

context

A page should mean "users are hurt or soon will be, and a human must act now". Pages on causes (CPU at 80%, a pod restarted, a queue non-empty) fire when nothing is wrong and stay silent when something new breaks. Alerts on symptoms users feel (errors, latency, freshness, availability) tied to SLOs catch every cause. Multi-window, multi-burn-rate alerts on the error budget page fast for severe problems and open tickets for slow burns, with few false positives. Everything else is a ticket, a dashboard, or deleted.

task

Design the alerts for:

service and metrics

Write rules in format.

  1. State the SLOs you will alert on. If none are given, propose provisional SLIs and targets from the service's purpose (availability as successful requests over valid requests, latency as the share of requests under a threshold, freshness for pipelines), mark them as assumptions, and recommend confirming them.
  2. Design burn-rate alerts per SLO. Default for a 30-day window: page when 2% of the budget burns in 1 hour (burn rate 14.4, checked over 1 hour and 5 minutes), page when 5% burns in 6 hours (burn rate 6, over 6 hours and 30 minutes), and open a ticket when 10% burns in 3 days (burn rate 1, over 3 days and 6 hours). Show the arithmetic for this service's target. Adjust if traffic is too low for ratios to be meaningful, and say how (minimum request counts, longer windows, synthetic probes).
  3. Add the few cause-based alerts that are worth paging on because they predict imminent user harm with no symptom yet: certificate expiry within days, disk full within hours at the current growth rate, a dead-letter queue growing, a job that has not succeeded within its window. Prefer predictive forms (time to full) over static thresholds.
  4. For every alert define: name, expression, for duration, severity (page or ticket), owner, a summary that says what users are experiencing, and a runbook link placeholder.
  5. Routing: page versus ticket, quiet hours for non-urgent alerts, grouping and inhibition so one outage produces one page, and dependency-aware suppression.
  6. Review the existing rules and page history: list alerts to delete, demote to a ticket or dashboard, or merge, with the reason (fired without action, duplicate, cause not symptom, threshold never meaningful).
constraints
  • Use only metric names and labels from the input; where you need one that is not there, write it as a placeholder and list it under Gaps.
  • Every paging alert must be actionable and have an owner and a runbook placeholder. If you cannot say what the responder would do, it does not page.
  • Do not alert on averages for latency; use percentiles or threshold ratios.
  • Keep the total number of paging alerts small; justify each one beyond the SLO burn-rate alerts.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Assumptions

Bullets, including provisional SLOs.

Alert design

A table: alert, type (burn-rate, predictive, cause), severity, why it pages or tickets, what the responder does.

Rules

One fenced block with all rules in the chosen format.

Routing

Bullets or a routing config sketch.

Delete or demote

A table: existing alert, action (delete, demote, merge), reason.

Gaps

Missing metrics or instrumentation needed, or "None".

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Incident and operations
level
Intermediate
made for
Site reliability engineer, DevOps / platform engineer, Backend engineer, Engineering manager
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install design-alerting-rules --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill design-alerting-rules -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PersonaIncident and operations

Site reliability engineer

Acts as a site reliability engineer who thinks in SLOs and error budgets, automates toil, designs for failure and writes blameless reviews.

site-reliability-engineer
PromptIncident and operations

Define SLOs and burn-rate alerts

Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.

define-slos
PromptIncident and operations

Write an operational runbook

Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.

write-runbook
PromptIncident and operations

Design an on-call rotation

Designs an on-call rotation with schedule, escalation, handoff, alert ownership, compensation norms and health checks. Use when starting on-call or when the current one burns people out.

design-on-call-rotation
PromptIncident and operations

Instrument a service for observability

Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.

instrument-service-observability
PromptIncident and operations

Plan a game day or chaos exercise

Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.

plan-game-day