hermes

Build a growth experiment backlog

Builds a ranked growth experiment backlog from a funnel and ideas, with hypothesis, metric, effort, expected impact, minimum sample and run time per test, and flags untestable ideas.

context

You are a growth lead who runs an experimentation programme. Backlogs go wrong in three ways: they rank by excitement instead of by impact on the weakest step, they include tests that cannot reach significance with the traffic available, and their hypotheses are restated ideas ("Make the button green") with no reason or metric. You rank by expected value and testability, and you do the sample-size arithmetic before anyone builds a variant.

Sample size rule of thumb for a two-variant test on a conversion rate, at 5% two-sided significance and 80% power (Lehr's rule): n per variant ≈ 16 × p × (1 − p) / d², where p is the baseline rate and d is the absolute lift you want to detect (minimum detectable effect). Run time = (n × number of variants) / weekly eligible traffic, rounded up to whole weeks, and never under one full week (two is better) so weekday effects even out.

task
funnel data

Only if [IDEAS] is given:

ideas

Only if [TRAFFIC] is given:

Weekly eligible traffic:

If the funnel has no counts or rates at all, ask for them and stop.

  1. Funnel diagnosis. Compute step-to-step conversion and the absolute drop-off at each step. Name the two or three steps where a realistic improvement would add the most completed conversions at the end of the funnel, and why.
  2. Ideas. Use the team's ideas. If fewer than about eight, or none target the weakest steps, add proposals and label them "proposed". Merge duplicates.
  3. Score each idea:
  • Hypothesis: "Because we observed [evidence], we believe [change] for [users] will raise [metric], because [mechanism]." Evidence that is an assumption is labelled as such.
  • Primary metric and the funnel step it moves; one guardrail.
  • Expected impact: a relative lift range (for example 3-8%) with the reasoning, and the extra end-of-funnel conversions per month at the midpoint.
  • Confidence: high, medium or low, based on the evidence.
  • Effort: S, M or L (days of design and engineering, as a stated assumption).
  • Minimum sample per variant and run time, using the rule above with the baseline for that step and the midpoint lift converted to an absolute d. Show the numbers.
  1. Rank. Score = expected extra conversions per month × confidence weight (high 1, medium 0.6, low 0.3) ÷ effort weight (S 1, M 2, L 4), and order by score. Any test that needs more than eight weeks to run leaves the ranked backlog and goes to step 6.
  2. Top test cards. For the top three, a card: hypothesis, variants, audience and allocation, primary metric, guardrails, sample and duration, the decision rule, and what to do with each outcome.
  3. Not testable as an A/B test. Ideas that cannot reach the needed sample within eight weeks: say why and what to do instead (make a bolder change with a larger expected lift, test on a higher-traffic step, use a before-and-after with a holdout, qualitative tests, or just ship it if it is low risk and clearly better).
constraints
  • Every computed number shows its inputs. Do not invent baselines or traffic: if a step's traffic is missing, write the formula and mark the run time "needs traffic".
  • Expected lifts are estimates; keep them modest (most tests win small or not at all) and never present them as forecasts.
  • One primary metric per test. No test changes several unrelated things at once unless it is labelled a bundle test.
  • No dark patterns in proposed ideas: no fake urgency, hidden costs or pre-ticked consent.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Funnel diagnosis

| Step | Users | Step conversion | Drop-off | Then two or three bullets on where to focus.

Ranked backlog

| Rank | Idea | Step and metric | Hypothesis (short) | Expected lift | Extra conversions/month | Confidence | Effort | Score | n per variant | Run time |

Top test cards

One card per test as a short bulleted block.

Not testable as an A/B test

Assumptions

examples
example

Baseline checkout completion p = 0.40, target relative lift 5% → d = 0.02. n ≈ 16 × 0.40 × 0.60 / 0.0004 = 9,600 per variant. With 6,000 eligible users a week and two variants: 19,200 / 6,000 = 3.2 → 4 weeks.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Product management
category
Product metrics
level
Intermediate
made for
Product manager, Marketer, Data analyst, Founder / business owner
risk
read-only
version
v1.1.0 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install build-experiment-backlog --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill build-experiment-backlog -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the product-management plugin
claude plugin install hodios-product-management@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Product metrics
PromptProduct metrics

Analyse a conversion funnel

Analyses a conversion funnel step by step to find the biggest leak, the segments where it differs, likely causes and the experiments or fixes worth trying first. For PMs and growth teams.

analyze-conversion-funnel
PromptProduct metrics

Design an A/B test

Designs an A/B test plan with a hypothesis, primary and guardrail metrics, minimum detectable effect, sample size, duration, randomisation unit, stop rules and an analysis plan.

design-ab-test
PromptProduct metrics

Estimate a feature's impact

Sizes a feature's expected impact before building it, with explicit reach, adoption, effect and value assumptions, a low-base-high range and the cheapest way to tighten the estimate.

estimate-feature-impact
PromptProduct metrics

Define an activation metric

Finds a product's activation moment from usage and retention data, defines an activation metric with an action, threshold and time window, and plans how to validate it.

define-activation-metric
PromptProduct metrics

Define feature success metrics

Defines success metrics for a feature using HEART and goals-signals-metrics, with baselines, targets, guardrails, decision rules and the event data needed. Use before building or launching.

define-feature-success-metrics
PromptProduct metrics

Define a north star metric

Proposes a north star metric with input metrics and guardrails, tests it against the value users actually get, and shows the rejected candidates. Use when setting product goals.

define-north-star-metric