Design an on-call rotation
Designs an on-call rotation with schedule, escalation, handoff, alert ownership, compensation norms and health checks. Use when starting on-call or when the current one burns people out.
On-call is sustainable when the rotation is big enough, the pages are few and actionable, handoffs carry context, and the people on it are compensated and rested. It fails when four people cover a week each with 30 pages a night, when nobody owns the noisy alerts, when the secondary is never paged so nobody knows if escalation works, or when time off after a bad night depends on asking. Common reference points: a primary and a secondary, at least six to eight people per around-the-clock rotation (or a follow-the-sun split across regions so nobody is paged at night), and a target of a few pages per shift at most, each one actionable.
Design on-call with coverage for:
- List what you know and what you assume: people, time zones, services and tiers, page volume, existing pay or policy. If headcount or page volume is missing, ask under Open questions and design with a stated assumption.
- Rotation. Pick the shape and justify it: weekly or split-week shifts, primary and secondary, follow-the-sun if there are two or more regions at least six hours apart. State the handover time (a working hour, mid-week rather than Monday or Friday), how often each person is on call per month, and the minimum headcount the shape needs. If the team is too small for the coverage, say so plainly and give options (reduce coverage tier for low-criticality services, share a rotation with another team, vendor support, business-hours only with best-effort nights).
- Escalation. Paging timeline: primary acknowledges within N minutes, then secondary, then the engineering manager or incident commander, with values per service tier. Include how to escalate to other teams and vendors, and when to declare an incident.
- Handoff. A short handoff template: open incidents, ongoing risks, noisy alerts, changes deployed, things to watch. Make the handoff synchronous for 10 to 15 minutes or written with acknowledgement.
- Alert ownership. Every paging alert has an owning team and a runbook link; anything without one does not page. The on-call engineer may silence a non-actionable alert and must file a ticket. Reserve on-call time for reliability work when it is quiet.
- Compensation and time off. Propose norms: pay or time-off-in-lieu per shift and per out-of-hours page, rest after a night page, no on-call in the first weeks for new joiners until they have shadowed. Tell the user to confirm with HR and local labour law, since rules differ by country.
- Health checks. Metrics to review monthly: pages per shift, out-of-hours pages, time to acknowledge, percentage of actionable pages, repeat alerts, and a short on-call survey. Set thresholds that trigger action (for example more than two out-of-hours pages per week).
- Rollout. Shadowing and reverse-shadowing, a paging test of the full escalation chain, and a review after the first month.
- Do not invent headcount, salaries, or legal requirements. Compensation is a proposal of norms with ranges or structures, not a figure for this company.
- Prefer fewer, actionable pages over more coverage; never solve noise by adding people.
- Keep it fair: the same rules apply to managers and senior engineers who are on the rotation.
- Times are written with a time zone. Where locations observe daylight saving on different dates, say how the shift boundaries move in those weeks.
Assumptions
Bullets.
Rotation
The shape, a table of shifts with times and who covers them (placeholders), and on-call frequency per person.
Escalation
A table by service tier: acknowledge target, escalate after, next level.
Handoff
The template in a fenced block.
Alert ownership
Rules as bullets.
Compensation and time off
Proposed norms, marked "confirm with HR and local law".
Health checks
A table: metric, target, action threshold.
Rollout
Numbered steps with dates or weeks.
Open questions
Numbered.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Incident and operations
- level
- Intermediate
- made for
- Engineering manager, Site reliability engineer, Tech lead / staff engineer, DevOps / platform engineer
- risk
- read-only
- version
- v1.0.1 · incubating
- reviewed
- 2026-10-03
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install design-on-call-rotation --target claude-codenpx skills add hermes-hq/hodios-dist --skill design-on-call-rotation -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Incident and operationsSite reliability engineer
Acts as a site reliability engineer who thinks in SLOs and error budgets, automates toil, designs for failure and writes blameless reviews.
site-reliability-engineerIncident commander
Runs a live incident like an experienced incident commander, assigning roles, keeping a steady comms cadence and driving mitigation before root cause. Use as the coordinating voice during an outage.
incident-commanderDesign actionable alerting rules
Designs actionable alerts from SLOs and user-facing symptoms, with thresholds, routing, runbook links, and a list of noisy alerts to delete. Use when pages are noisy or real outages go unnoticed.
design-alerting-rulesWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookDefine SLOs and burn-rate alerts
Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.
define-slosInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observability