Plan backups and disaster recovery
Writes a backup and disaster-recovery plan with RPO and RTO targets, dependency order, restore drills and owner checklists. Use when a system has backups nobody has restored, or no plan at all.
Disaster-recovery plans fail on the things nobody listed: backups that were never restored, replicas that faithfully copied the corruption, the secrets manager or the backup credentials living in the region that went down, a DNS change only one person knows how to make. A useful plan is specific to the system, measured in minutes and data lost, and proven by drills.
Write a backup and disaster-recovery plan for: Recovery point objective: Only if [RPO] is given: Recovery time objective: Only if [RTO] is given:
- If an objective above is blank, propose per-tier targets with reasoning and mark them "proposed, needs business sign-off". Do not present them as decided.
- Inventory every component and data store. Assign each a tier, and list what it depends on to start: identity, secrets, DNS, certificates, container registry, CI/CD, third-party APIs.
- Cover these scenarios separately, because each needs a different answer: accidental deletion, logical corruption (replication copies it, so point-in-time recovery is required), loss of a zone, loss of a region, compromised cloud account or ransomware, and a critical vendor outage.
- For each data store, specify the backup method, frequency (it must meet the RPO), retention, encryption and where the key lives, and isolation: a separate account or immutable storage so an attacker with production access cannot delete backups.
- Choose a recovery strategy per tier (backup and restore, pilot light, warm standby or active-active) and justify it against the RTO and cost.
- Write the recovery order from the dependency graph: what must be up before what, with an estimated time per step and a total compared against the RTO.
- Define restore drills: what is restored, how often, success criteria (measured RPO and RTO), and who signs off.
- Replication and high availability are not backups. Do not count them toward recovery from corruption or deletion.
- A backup is only counted as working once a restore of it has been tested. Mark untested backups as risks.
- Use the details given. Where a fact is missing (sizes, regions, owners), write a clearly marked placeholder and list it under Open risks rather than inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Summary
The targets (stated or proposed), the strategy per tier, and the three biggest gaps today.
Inventory
A table: component, tier, data store (yes/no), depends on, current backup, gap.
Scenarios
One short subsection per scenario: detection, decision owner, recovery path, expected data loss and downtime.
Backup policy
A table: data store, method, frequency, retention, isolation, encryption key location, last tested restore.
Recovery order
Numbered steps with estimated durations and a total against the RTO.
Drills
A table: drill, frequency, success criteria, owner.
Owner checklists
One checklist per role (for example incident lead, database owner, platform owner).
Open risks
Bullets: missing information and unproven assumptions.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- DevOps
- level
- Expert
- made for
- Site reliability engineer, DevOps / platform engineer, Software architect, Engineering manager
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install plan-disaster-recovery --target claude-codenpx skills add hermes-hq/hodios-dist --skill plan-disaster-recovery -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of DevOpsWrite an operational runbook
Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.
write-runbookDesign a deployment strategy
Chooses and specifies a deployment strategy (rolling, blue-green, canary or feature-flagged) with health gates, automated rollback triggers and database-change ordering. Use when deploys feel risky.
design-deployment-strategyReview a Dockerfile
Reviews a Dockerfile for security, image size, build cache use and runtime correctness, and returns ranked findings with a corrected file. Use before shipping a new or changed container image.
review-dockerfileReview an infrastructure plan before apply
Reviews a Terraform, OpenTofu or other IaC plan for destructive changes, security exposure, cost surprises and changes outside the stated intent. Use before running apply, especially in production.
review-iac-planWrite a Docker Compose dev environment
Writes a Docker Compose local development setup that mirrors production dependencies, with health checks, named volumes, seed data, env files and a one-command start. Use when onboarding developers.
write-docker-composeWrite a GitHub Actions workflow
Writes a secure, cached and least-privilege GitHub Actions workflow that fits the repository's real build and test commands. Use when adding CI, a release job or a scheduled task.
write-github-actions-workflow