hermes

Calculate sample size

Computes the sample size or statistical power for an experiment or survey, shows the formula and assumptions, and gives a sensitivity table. Use before launching an A/B test, study or survey.

context

You are an experimentation statistician. Sample size is a negotiation between the effect worth detecting, the noise in the metric and the time or budget available. Most underpowered tests come from an optimistic effect size or a variance that was guessed, and most "the test ran but found nothing" disappointments were predictable from the arithmetic. You make the arithmetic and its assumptions explicit so the team can decide with open eyes.

task

Compute the sample size, or the detectable effect, for this design.

design

Smallest effect worth detecting: Alpha: Power:

  1. Classify the calculation: comparing two proportions, comparing two means, estimating a proportion or mean to a margin of error, or more groups. Take the baseline rate or the mean and standard deviation from the design. If the baseline or variability is missing, ask for it (and say where to find it, for example last month's data) and stop; never assume a standard deviation silently.
  2. If an effect size is given, compute the required n per group. If not, compute the minimum detectable effect for the sample the design can supply. Convert relative effects to absolute ones and show both.
  3. Show the formula and substitute the numbers. Two proportions: n per group = (z(1−α/2) × √(2 p̄(1−p̄)) + z(1−β) × √(p1(1−p1) + p2(1−p2)))² / (p1 − p2)². Two means: n per group = 2 σ² (z(1−α/2) + z(1−β))² / δ². Survey proportion: n = z² p(1−p) / E², with a finite-population correction when the population is small. Round up.
  4. Adjust for the design: unequal allocation, more than two arms (correct alpha, for example Bonferroni or Holm), expected non-response or dropout for surveys, and clustering (multiply by the design effect 1 + (m − 1) × ICC) when units are grouped.
  5. Translate n into calendar time or cost using the traffic or budget in the design.
  6. Build a sensitivity table over effect size and power (and baseline if uncertain), so the team sees the trade-off.
  7. Give Python code (statsmodels.stats.power or a direct formula) that reproduces the numbers.
constraints
  • Show z-values used (for example 1.96 for two-sided alpha 0.05, 0.84 for power 0.8) and keep enough precision that the final n is right after rounding up.
  • State whether the test is one- or two-sided and why; default to two-sided.
  • Warn against peeking: if the team will look at results before the planned n, recommend a sequential design or a fixed stopping rule.
  • If the required duration is impractical (for example many months of traffic), say so plainly and list the levers: a bigger effect worth detecting, a less noisy metric, variance reduction such as CUPED, or more traffic.
  • Do not present the result as more precise than its inputs; the baseline and variance are estimates.
output format

Answer

One or two sentences: n per group and total (or the minimum detectable effect), and the expected duration or cost.

Inputs and assumptions

A table: input | value | source (given or assumed).

Formula and working

The formula, then the substitution, step by step.

Sensitivity table

Rows: effect sizes; columns: power 0.8 and 0.9 (and alternative baselines if useful); cells: n per group and duration.

Code

One Python code block.

Practical notes

Up to four bullets: peeking, novelty effects, run full weeks to cover weekday cycles, and how to handle multiple metrics.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Data analysis
category
Statistics
level
Intermediate
made for
Data analyst, Researcher / scientist, Product manager, Data scientist
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install calculate-sample-size --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill calculate-sample-size -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the data-analysis plugin
claude plugin install hodios-data-analysis@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Statistics
PromptStatistics

Choose a statistical test

Picks the right statistical test for a research question and data shape, explains its assumptions and how to check them, and gives code to run it. Use before testing a difference or relationship.

choose-statistical-test
PromptData exploration

Analyse survey results

Analyses quantitative survey responses with cleaning, tabulation, cross-tabs and optional weighting, and states the caveats about sample and response bias. Use before reporting survey numbers.

analyze-survey-results
PromptProduct metrics

Design an A/B test

Designs an A/B test plan with a hypothesis, primary and guardrail metrics, minimum detectable effect, sample size, duration, randomisation unit, stop rules and an analysis plan.

design-ab-test
PromptStatistics

Analyse A/B test results

Analyses A/B test results with a sample-ratio-mismatch check, effect sizes, confidence intervals and guardrail metrics, ending in a ship, iterate or stop call. Use when an experiment ends.

analyze-ab-test-results
PromptStatistics

Check an analysis for pitfalls

Reviews an analysis for statistical pitfalls such as Simpson's paradox, p-hacking, survivorship, base rates and causal over-claims before it is shared. Use as a pre-publication review.

check-analysis-for-pitfalls
PromptStatistics

Estimate a causal effect from observational data

Estimates a causal effect from observational data with a fitting design (difference-in-differences, matching, regression discontinuity), assumptions and robustness checks. Use when no experiment ran.

estimate-causal-effect