hermes

Design a deployment strategy

Chooses and specifies a deployment strategy (rolling, blue-green, canary or feature-flagged) with health gates, automated rollback triggers and database-change ordering. Use when deploys feel risky.

context

Deploys are scary when a bad release reaches every user at once, when nobody knows it is bad until customers complain, and when rollback is a manual procedure that has never been practised. The fix is not one technique but a combination sized to the system: limiting how many users see a release before it is trusted, automated checks that compare the new version against the old, a rollback that is one action and tested, schema changes ordered so old and new code both work, and separating deploying code from releasing features. Each technique has costs: blue-green needs double capacity, a canary needs enough traffic to produce a signal, and feature flags add code paths that must be cleaned up.

task

Design the deployment strategy for: Only if [PLATFORM] is given:

Platform: Only if [RISK_PROFILE] is given:

Risk profile: Only if [CURRENT_PROCESS] is given:

Current process:

  1. Choose the strategy and justify it against the system's properties: stateless or stateful, traffic volume (enough requests in a canary slice to detect a regression within minutes), long-lived connections or sessions, client versions you do not control (mobile apps, partner integrations), capacity cost, and the risk profile. Say why the alternatives are worse here. Combine techniques where it helps, for example a canary for the deploy plus feature flags for risky behaviour changes.
  2. Define the rollout stages: traffic share or instance count per stage, bake time per stage, and whether each promotion is automatic or needs approval.
  3. Define health gates for each stage: pre-traffic checks (readiness, smoke tests against the new version), and live comparisons of the new version against the current one on error rate, latency percentiles, saturation and one business signal (checkouts, sign-ins). Give each gate a threshold, a comparison window and a minimum sample size, as starting values.
  4. Define automated rollback triggers: which gate failures roll back without a human, how fast, and what alerts and records are produced. Say which failures should page someone even after an automatic rollback.
  5. Order database and schema changes with expand and contract: migrations must work with both the current and the new code; destructive steps ship in a later release after the old code is gone; backfills run separately and are throttled. State the rule for what may ship together in one deploy.
  6. Specify the rollback procedure: one command or button, how long it takes, what it does not undo (migrations, messages already sent, cache entries, data written in a new format), and how often it is rehearsed.
  7. Describe the implementation on the platform: which native features or tools provide traffic splitting, analysis and rollback, the pipeline stages, deploy markers on dashboards, and deploy freeze rules.
  8. Plan the move from the current process in small steps, each one an improvement on its own.

If the system description lacks traffic volume, state handling or how rollback works today, and the choice depends on it, ask for it before choosing. Otherwise state assumptions.

constraints
  • Choose the simplest strategy that meets the risk profile. A low-traffic internal tool does not need a five-stage canary.
  • Every threshold is a starting value with the reason for it, to be tuned from real deploys.
  • Name tools only as examples of a capability available on the platform.
  • Never treat "roll back" as free: list what a rollback cannot undo.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Recommendation

The strategy in two or three sentences, and why the alternatives lose.

Rollout stages

Table: stage | traffic or instances | bake time | promotion (automatic or approval).

Health gates and rollback triggers

Table: signal | comparison | threshold | window | action on failure.

Database and schema changes

Numbered rules, then an example sequence for a column rename across releases.

Rollback procedure

Steps, expected duration, and what it does not undo.

Implementation

How to build it on the platform, with a pipeline sketch as a code block in the platform's format where possible.

Migration plan

Numbered steps from today's process to the target.

Risks and open questions

Bullets.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
DevOps
level
Intermediate
made for
DevOps / platform engineer, Site reliability engineer, Tech lead / staff engineer, Backend engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install design-deployment-strategy --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill design-deployment-strategy -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of DevOps
PromptImplementation

Put a change behind a feature flag

Wraps new behaviour behind a feature flag with a safe default, a kill switch, tests for both paths and a cleanup ticket. Use when shipping a risky change incrementally.

add-feature-flag
PromptData engineering

Plan a zero-downtime schema change

Turns current table DDL and a desired change into expand and contract steps with lock-safe SQL, app changes, backfill, verification and rollback. Use before altering a live table.

plan-zero-downtime-schema-change
PromptIncident and operations

Define SLOs and burn-rate alerts

Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.

define-slos
PromptDevOps

Write a GitHub Actions workflow

Writes a secure, cached and least-privilege GitHub Actions workflow that fits the repository's real build and test commands. Use when adding CI, a release job or a scheduled task.

write-github-actions-workflow
PersonaDevOps

DevOps engineer

Acts as a DevOps engineer who automates the second time, keeps pipelines fast and reproducible, and makes every change reversible. Use for CI/CD, infrastructure and release work.

devops-engineer
PromptDevOps

Review a Dockerfile

Reviews a Dockerfile for security, image size, build cache use and runtime correctness, and returns ranked findings with a corrected file. Use before shipping a new or changed container image.

review-dockerfile