hermes

Design an on-call rotation

Designs an on-call rotation with schedule, escalation, handoff, alert ownership, compensation norms and health checks. Use when starting on-call or when the current one burns people out.

context

On-call is sustainable when the rotation is big enough, the pages are few and actionable, handoffs carry context, and the people on it are compensated and rested. It fails when four people cover a week each with 30 pages a night, when nobody owns the noisy alerts, when the secondary is never paged so nobody knows if escalation works, or when time off after a bad night depends on asking. Common reference points: a primary and a secondary, at least six to eight people per around-the-clock rotation (or a follow-the-sun split across regions so nobody is paged at night), and a target of a few pages per shift at most, each one actionable.

task

Design on-call with coverage for:

team and services

  1. List what you know and what you assume: people, time zones, services and tiers, page volume, existing pay or policy. If headcount or page volume is missing, ask under Open questions and design with a stated assumption.
  2. Rotation. Pick the shape and justify it: weekly or split-week shifts, primary and secondary, follow-the-sun if there are two or more regions at least six hours apart. State the handover time (a working hour, mid-week rather than Monday or Friday), how often each person is on call per month, and the minimum headcount the shape needs. If the team is too small for the coverage, say so plainly and give options (reduce coverage tier for low-criticality services, share a rotation with another team, vendor support, business-hours only with best-effort nights).
  3. Escalation. Paging timeline: primary acknowledges within N minutes, then secondary, then the engineering manager or incident commander, with values per service tier. Include how to escalate to other teams and vendors, and when to declare an incident.
  4. Handoff. A short handoff template: open incidents, ongoing risks, noisy alerts, changes deployed, things to watch. Make the handoff synchronous for 10 to 15 minutes or written with acknowledgement.
  5. Alert ownership. Every paging alert has an owning team and a runbook link; anything without one does not page. The on-call engineer may silence a non-actionable alert and must file a ticket. Reserve on-call time for reliability work when it is quiet.
  6. Compensation and time off. Propose norms: pay or time-off-in-lieu per shift and per out-of-hours page, rest after a night page, no on-call in the first weeks for new joiners until they have shadowed. Tell the user to confirm with HR and local labour law, since rules differ by country.
  7. Health checks. Metrics to review monthly: pages per shift, out-of-hours pages, time to acknowledge, percentage of actionable pages, repeat alerts, and a short on-call survey. Set thresholds that trigger action (for example more than two out-of-hours pages per week).
  8. Rollout. Shadowing and reverse-shadowing, a paging test of the full escalation chain, and a review after the first month.
constraints
  • Do not invent headcount, salaries, or legal requirements. Compensation is a proposal of norms with ranges or structures, not a figure for this company.
  • Prefer fewer, actionable pages over more coverage; never solve noise by adding people.
  • Keep it fair: the same rules apply to managers and senior engineers who are on the rotation.
  • Times are written with a time zone. Where locations observe daylight saving on different dates, say how the shift boundaries move in those weeks.
output format

Assumptions

Bullets.

Rotation

The shape, a table of shifts with times and who covers them (placeholders), and on-call frequency per person.

Escalation

A table by service tier: acknowledge target, escalate after, next level.

Handoff

The template in a fenced block.

Alert ownership

Rules as bullets.

Compensation and time off

Proposed norms, marked "confirm with HR and local law".

Health checks

A table: metric, target, action threshold.

Rollout

Numbered steps with dates or weeks.

Open questions

Numbered.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Incident and operations
level
Intermediate
made for
Engineering manager, Site reliability engineer, Tech lead / staff engineer, DevOps / platform engineer
risk
read-only
version
v1.0.1 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install design-on-call-rotation --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill design-on-call-rotation -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PersonaIncident and operations

Site reliability engineer

Acts as a site reliability engineer who thinks in SLOs and error budgets, automates toil, designs for failure and writes blameless reviews.

site-reliability-engineer
PersonaIncident and operations

Incident commander

Runs a live incident like an experienced incident commander, assigning roles, keeping a steady comms cadence and driving mitigation before root cause. Use as the coordinating voice during an outage.

incident-commander
PromptIncident and operations

Design actionable alerting rules

Designs actionable alerts from SLOs and user-facing symptoms, with thresholds, routing, runbook links, and a list of noisy alerts to delete. Use when pages are noisy or real outages go unnoticed.

design-alerting-rules
PromptIncident and operations

Write an operational runbook

Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.

write-runbook
PromptIncident and operations

Define SLOs and burn-rate alerts

Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.

define-slos
PromptIncident and operations

Instrument a service for observability

Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.

instrument-service-observability