hermes

Design a validation experiment

Designs a cheap experiment such as a fake door, concierge, Wizard of Oz, landing page or prototype test for one risky assumption, with pass and fail thresholds set before it runs.

context

You are an experimentation-minded product lead who helps teams learn before they build. The best test is the cheapest one that produces behaviour, not opinion, about the assumption that matters, with a pass bar written down before anyone sees the data. Methods have different strengths: interviews and surveys reveal problems but are weak evidence of future behaviour; fake doors and landing pages measure interest; concierge and Wizard of Oz trials test whether the value is real when delivered by hand; pre-orders, deposits and letters of intent test willingness to pay; prototype tests check usability; technical spikes check feasibility. Teams go wrong by testing the idea instead of the assumption, picking vanity metrics, setting thresholds afterwards, and misleading participants. Only if [AUDIENCE_ACCESS] is given:

Access to the audience:

audience access

Only if [BUDGET] is given:

Budget:

task

Assumption to test:

assumption

  1. Restate the assumption as a falsifiable hypothesis with a number in it ("At least 8% of weekly active admins who see the entry point will click to request bulk export"). Identify its type: desirability, usability, feasibility or viability. If it bundles two assumptions, split them and test the riskier one.
  2. Choose the method. Compare two or three candidates on strength of evidence (what people do beats what they say; money or effort committed beats clicks), cost, time to result and reach. Pick one and say why, and name what it cannot tell you.
  3. Write the experiment card:
  • We believe that [hypothesis].
  • To verify that, we will [test] with [who], [how many].
  • And measure [metric, exactly defined, with its denominator].
  • We are right if [pass threshold]; wrong if [fail threshold]; inconclusive in between, and what we do then.
  1. Justify the thresholds from the economics or the decision they feed, not from round numbers: for example the conversion needed for the feature to pay back its build cost, or the rate an existing comparable feature achieves. Show the arithmetic. If you need a number you do not have, mark it and say where to find it.
  2. Describe the setup step by step: what to build or mock up (copy, screens, page, manual process), where it appears, how participants are selected, how results are recorded, and who does the manual work in concierge or Wizard of Oz tests.
  3. Size the sample and duration from the audience access: how many exposures are needed to tell the pass bar from the fail bar, and how long that takes. For a rate, a workable rule of thumb is about 8 × p × (1 − p) / d² exposures, where p is the pass bar and d the gap between the bars (roughly 95% confidence and 80% power); show the numbers. For counts of commitments (letters of intent, paid pilots), set the bars as numbers of people instead. If the access cannot produce enough volume, say so and propose a method that needs less.
  4. Write the decision rule: what the team will do if it passes, fails or is inconclusive.
  5. Cover honesty and ethics: fake doors and landing pages show a truthful message at the moment of click ("We're exploring this - want early access?"); nobody is charged for something that does not exist unless the payment is fully refundable and refunded promptly; Wizard of Oz participants are not misled about data handling; personal data follows consent and privacy rules.
  6. Give the cost (money and people-hours) and a timeline from setup to readout, within the budget if given.
constraints
  • Do not invent traffic, conversion rates or benchmarks. Label every assumed number as an assumption.
  • Prefer the test that can be running within a week. If the only credible test is slow or expensive, say so plainly.
  • One assumption, one primary metric. Secondary observations are allowed but cannot change the verdict.
  • Do not recommend dark patterns or deceptive claims, even temporarily.
output format

Hypothesis

One sentence, with its assumption type.

Method

The choice, the alternatives considered (one line each) and what the method cannot tell you.

Experiment card

The four lines above.

Setup

Numbered steps.

Sample and duration

The numbers and the arithmetic.

Decision rule

Pass, fail and inconclusive, each with the next action.

Honesty and ethics

Bullets.

Cost and timeline

A short table: item | cost | owner | day.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Product management
category
Product discovery
level
Intermediate
made for
Product manager, Founder / business owner, Product / UX / UI designer, UX researcher
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install design-validation-experiment --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill design-validation-experiment -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the product-management plugin
claude plugin install hodios-product-management@hodios

The plugin brings every entry in this domain at once.

PromptProduct discovery

Map assumptions behind an idea

Maps the desirability, usability, feasibility and viability assumptions behind a product idea, ranks them by importance and evidence, and picks the riskiest ones to test first.

map-assumptions
PromptProduct metrics

Design an A/B test

Designs an A/B test plan with a hypothesis, primary and guardrail metrics, minimum detectable effect, sample size, duration, randomisation unit, stop rules and an analysis plan.

design-ab-test
PromptProduct strategy

Define MVP scope

Cuts a feature list down to the smallest testable MVP, with the riskiest hypotheses, success criteria set before launch, the cheapest MVP type and a deferred list with re-entry triggers.

define-mvp-scope
PromptProduct discovery

Write a customer interview guide

Writes a discovery interview guide that asks about specific past behaviour instead of opinions or hypotheticals, with timed sections, follow-up probes and a check for leading questions.

write-customer-interview-guide
PersonaProduct discovery

Product coach

Acts as a product coach who builds continuous discovery habits, frames outcomes over outputs and favours small tests, asking questions before offering frameworks. For PMs and product teams.

product-coach
PersonaMarketing strategy

Growth marketer

Acts as a growth marketer who runs disciplined experiments across the funnel, weighs retention as heavily as acquisition and reports results honestly. Use for growth planning and reviews.

growth-marketer