hermes

Define SLOs and burn-rate alerts

Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.

context

Teams write SLOs that measure servers instead of users ("CPU below 80%"), pick 99.99% because it sounds good, and alert on raw error rate, which pages for blips and misses slow burns. A good SLO measures what users experience on a journey, sets a target the service can meet and users would accept, and alerts on how fast the error budget is burning.

task

Define SLOs for from these user journeys: Only if [CURRENT_METRICS] is given: Current metrics and performance:

  1. For each journey, choose 1 or 2 SLIs written as good events divided by valid events: availability, latency below a threshold, freshness or correctness. Say where each is measured (load balancer, server, client) and the trade-off. Define valid events explicitly, for example excluding health checks and client errors the user caused.
  2. Set a target and a window (a 28- or 30-day rolling window by default). Base the target on current performance and user need. If current metrics are missing, mark targets "provisional" and propose a 2 to 4 week baseline measurement.
  3. Compute the error budget in allowed bad events and in minutes of full outage per window.
  4. Write an error-budget policy: what happens at 50%, 75% and 100% consumed (for example: slow down risky launches, prioritise reliability work, freeze non-critical changes), the exceptions, and who decides.
  5. Write multi-window, multi-burn-rate alerts for a 30-day window: page at 14.4x burn over 1 hour (with a 5-minute short window), page at 6x over 6 hours (30-minute short window), and open a ticket at 1x over 3 days (6-hour short window). Adjust the numbers if the window differs and show the calculation.
  6. Write the alert rules in the syntax of the user's monitoring stack (PromQL recording and alerting rules by default). Note the low-traffic problem and a mitigation if any journey has little traffic.
constraints
  • No target of 100%, and no target tighter than the service's dependencies allow without saying how.
  • Use the metric names given; where none are given, use clearly named placeholders and say so.
  • Prefer few SLOs that matter over full coverage. Three per service is often enough.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

SLOs

A table: journey, SLI (good / valid), measured at, target, window, error budget.

Rationale

One short paragraph per SLO: why this SLI and target.

Error-budget policy

Thresholds, actions, exceptions, decision owner.

Alert rules

Fenced code blocks with the rules, then a table: alert, burn rate, long window, short window, budget consumed when it fires, page or ticket.

Open questions

What to confirm with product owners or measure first.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Incident and operations
level
Expert
made for
Site reliability engineer, DevOps / platform engineer, Engineering manager, Backend engineer
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install define-slos --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill define-slos -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptIncident and operations

Write an operational runbook

Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.

write-runbook
PromptIncident and operations

Instrument a service for observability

Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.

instrument-service-observability
PromptIncident and operations

Plan a game day or chaos exercise

Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.

plan-game-day
PromptIncident and operations

Triage a production alert

Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.

triage-production-alert
PromptIncident and operations

Write an incident status update

Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.

write-incident-update
PromptIncident and operations

Write a blameless postmortem

Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.

write-postmortem