hermes

Plan backups and disaster recovery

Writes a backup and disaster-recovery plan with RPO and RTO targets, dependency order, restore drills and owner checklists. Use when a system has backups nobody has restored, or no plan at all.

context

Disaster-recovery plans fail on the things nobody listed: backups that were never restored, replicas that faithfully copied the corruption, the secrets manager or the backup credentials living in the region that went down, a DNS change only one person knows how to make. A useful plan is specific to the system, measured in minutes and data lost, and proven by drills.

task

Write a backup and disaster-recovery plan for: Recovery point objective: Only if [RPO] is given: Recovery time objective: Only if [RTO] is given:

  1. If an objective above is blank, propose per-tier targets with reasoning and mark them "proposed, needs business sign-off". Do not present them as decided.
  2. Inventory every component and data store. Assign each a tier, and list what it depends on to start: identity, secrets, DNS, certificates, container registry, CI/CD, third-party APIs.
  3. Cover these scenarios separately, because each needs a different answer: accidental deletion, logical corruption (replication copies it, so point-in-time recovery is required), loss of a zone, loss of a region, compromised cloud account or ransomware, and a critical vendor outage.
  4. For each data store, specify the backup method, frequency (it must meet the RPO), retention, encryption and where the key lives, and isolation: a separate account or immutable storage so an attacker with production access cannot delete backups.
  5. Choose a recovery strategy per tier (backup and restore, pilot light, warm standby or active-active) and justify it against the RTO and cost.
  6. Write the recovery order from the dependency graph: what must be up before what, with an estimated time per step and a total compared against the RTO.
  7. Define restore drills: what is restored, how often, success criteria (measured RPO and RTO), and who signs off.
constraints
  • Replication and high availability are not backups. Do not count them toward recovery from corruption or deletion.
  • A backup is only counted as working once a restore of it has been tested. Mark untested backups as risks.
  • Use the details given. Where a fact is missing (sizes, regions, owners), write a clearly marked placeholder and list it under Open risks rather than inventing it.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Summary

The targets (stated or proposed), the strategy per tier, and the three biggest gaps today.

Inventory

A table: component, tier, data store (yes/no), depends on, current backup, gap.

Scenarios

One short subsection per scenario: detection, decision owner, recovery path, expected data loss and downtime.

Backup policy

A table: data store, method, frequency, retention, isolation, encryption key location, last tested restore.

Recovery order

Numbered steps with estimated durations and a total against the RTO.

Drills

A table: drill, frequency, success criteria, owner.

Owner checklists

One checklist per role (for example incident lead, database owner, platform owner).

Open risks

Bullets: missing information and unproven assumptions.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
DevOps
level
Expert
made for
Site reliability engineer, DevOps / platform engineer, Software architect, Engineering manager
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install plan-disaster-recovery --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill plan-disaster-recovery -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of DevOps
PromptIncident and operations

Write an operational runbook

Writes a runbook for an alert or routine procedure with symptoms, diagnosis commands, ordered mitigations, verification and escalation. Use so on-call engineers can act without tribal knowledge.

write-runbook
PromptDevOps

Design a deployment strategy

Chooses and specifies a deployment strategy (rolling, blue-green, canary or feature-flagged) with health gates, automated rollback triggers and database-change ordering. Use when deploys feel risky.

design-deployment-strategy
PromptDevOps

Review a Dockerfile

Reviews a Dockerfile for security, image size, build cache use and runtime correctness, and returns ranked findings with a corrected file. Use before shipping a new or changed container image.

review-dockerfile
PromptDevOps

Review an infrastructure plan before apply

Reviews a Terraform, OpenTofu or other IaC plan for destructive changes, security exposure, cost surprises and changes outside the stated intent. Use before running apply, especially in production.

review-iac-plan
PromptDevOps

Write a Docker Compose dev environment

Writes a Docker Compose local development setup that mirrors production dependencies, with health checks, named volumes, seed data, env files and a one-command start. Use when onboarding developers.

write-docker-compose
PromptDevOps

Write a GitHub Actions workflow

Writes a secure, cached and least-privilege GitHub Actions workflow that fits the repository's real build and test commands. Use when adding CI, a release job or a scheduled task.

write-github-actions-workflow