Design actionable alerting rules
Designs actionable alerts from SLOs and user-facing symptoms, with thresholds, routing, runbook links, and a list of noisy alerts to delete. Use when pages are noisy or real outages go unnoticed.
A page should mean "users are hurt or soon will be, and a human must act now". Pages on causes (CPU at 80%, a pod restarted, a queue non-empty) fire when nothing is wrong and stay silent when something new breaks. Alerts on symptoms users feel (errors, latency, freshness, availability) tied to SLOs catch every cause. Multi-window, multi-burn-rate alerts on the error budget page fast for severe problems and open tickets for slow burns, with few false positives. Everything else is a ticket, a dashboard, or deleted.
Design the alerts for:
Write rules in format.
- State the SLOs you will alert on. If none are given, propose provisional SLIs and targets from the service's purpose (availability as successful requests over valid requests, latency as the share of requests under a threshold, freshness for pipelines), mark them as assumptions, and recommend confirming them.
- Design burn-rate alerts per SLO. Default for a 30-day window: page when 2% of the budget burns in 1 hour (burn rate 14.4, checked over 1 hour and 5 minutes), page when 5% burns in 6 hours (burn rate 6, over 6 hours and 30 minutes), and open a ticket when 10% burns in 3 days (burn rate 1, over 3 days and 6 hours). Show the arithmetic for this service's target. Adjust if traffic is too low for ratios to be meaningful, and say how (minimum request counts, longer windows, synthetic probes).
- Add the few cause-based alerts that are worth paging on because they predict imminent user harm with no symptom yet: certificate expiry within days, disk full within hours at the current growth rate, a dead-letter queue growing, a job that has not succeeded within its window. Prefer predictive forms (time to full) over static thresholds.
- For every alert define: name, expression,
forduration, severity (page or ticket), owner, a summary that says what users are experiencing, and a runbook link placeholder. - Routing: page versus ticket, quiet hours for non-urgent alerts, grouping and inhibition so one outage produces one page, and dependency-aware suppression.
- Review the existing rules and page history: list alerts to delete, demote to a ticket or dashboard, or merge, with the reason (fired without action, duplicate, cause not symptom, threshold never meaningful).
- Use only metric names and labels from the input; where you need one that is not there, write it as a placeholder and list it under Gaps.
- Every paging alert must be actionable and have an owner and a runbook placeholder. If you cannot say what the responder would do, it does not page.
- Do not alert on averages for latency; use percentiles or threshold ratios.
- Keep the total number of paging alerts small; justify each one beyond the SLO burn-rate alerts.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Assumptions
Bullets, including provisional SLOs.
Alert design
A table: alert, type (burn-rate, predictive, cause), severity, why it pages or tickets, what the responder does.
Rules
One fenced block with all rules in the chosen format.
Routing
Bullets or a routing config sketch.
Delete or demote
A table: existing alert, action (delete, demote, merge), reason.
Gaps
Missing metrics or instrumentation needed, or "None".
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Intermediate
- made for
- Site reliability engineer, DevOps / platform engineer, Backend engineer, Engineering manager
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-03
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install design-alerting-rules --target claude-codenpx skills add hermes-hq/hodios-dist --skill design-alerting-rules -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsSite reliability engineer
Acts as a site reliability engineer who thinks in SLOs and error budgets, automates toil, designs for failure and writes blameless reviews.
site-reliability-engineerDefine SLOs and burn-rate alerts
Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.
define-slosWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookDesign an on-call rotation
Designs an on-call rotation with schedule, escalation, handoff, alert ownership, compensation norms and health checks. Use when starting on-call or when the current one burns people out.
design-on-call-rotationInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observabilityPlan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
plan-game-day