Define SLOs and burn-rate alerts
Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.
Teams write SLOs that measure servers instead of users ("CPU below 80%"), pick 99.99% because it sounds good, and alert on raw error rate, which pages for blips and misses slow burns. A good SLO measures what users experience on a journey, sets a target the service can meet and users would accept, and alerts on how fast the error budget is burning.
Define SLOs for from these user journeys: Only if [CURRENT_METRICS] is given: Current metrics and performance:
- For each journey, choose 1 or 2 SLIs written as good events divided by valid events: availability, latency below a threshold, freshness or correctness. Say where each is measured (load balancer, server, client) and the trade-off. Define valid events explicitly, for example excluding health checks and client errors the user caused.
- Set a target and a window (a 28- or 30-day rolling window by default). Base the target on current performance and user need. If current metrics are missing, mark targets "provisional" and propose a 2 to 4 week baseline measurement.
- Compute the error budget in allowed bad events and in minutes of full outage per window.
- Write an error-budget policy: what happens at 50%, 75% and 100% consumed (for example: slow down risky launches, prioritise reliability work, freeze non-critical changes), the exceptions, and who decides.
- Write multi-window, multi-burn-rate alerts for a 30-day window: page at 14.4x burn over 1 hour (with a 5-minute short window), page at 6x over 6 hours (30-minute short window), and open a ticket at 1x over 3 days (6-hour short window). Adjust the numbers if the window differs and show the calculation.
- Write the alert rules in the syntax of the user's monitoring stack (PromQL recording and alerting rules by default). Note the low-traffic problem and a mitigation if any journey has little traffic.
- No target of 100%, and no target tighter than the service's dependencies allow without saying how.
- Use the metric names given; where none are given, use clearly named placeholders and say so.
- Prefer few SLOs that matter over full coverage. Three per service is often enough.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
SLOs
A table: journey, SLI (good / valid), measured at, target, window, error budget.
Rationale
One short paragraph per SLO: why this SLI and target.
Error-budget policy
Thresholds, actions, exceptions, decision owner.
Alert rules
Fenced code blocks with the rules, then a table: alert, burn rate, long window, short window, budget consumed when it fires, page or ticket.
Open questions
What to confirm with product owners or measure first.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Expert
- made for
- Site reliability engineer, DevOps / platform engineer, Engineering manager, Backend engineer
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install define-slos --target claude-codenpx skills add hermes-hq/hodios-dist --skill define-slos -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observabilityPlan a game day or chaos exercise
Plans a game day or chaos exercise with failure scenarios, hypotheses, blast-radius limits, abort criteria, roles, an observation checklist and a follow-up review. Use to test resilience.
plan-game-dayTriage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
triage-production-alertWrite an incident status update
Writes a clear status update for an ongoing incident, tuned to customers, internal teams or executives, without speculation or promises the team cannot keep. Use for status pages, Slack and email.
write-incident-updateWrite a blameless postmortem
Turns incident notes, chat logs and timelines into a blameless postmortem with impact, timeline, contributing factors and owned action items. Use after an incident is resolved.
write-postmortem