Analyse A/B test results
Analyses A/B test results with a sample-ratio-mismatch check, effect sizes, confidence intervals and guardrail metrics, ending in a ship, iterate or stop call. Use when an experiment ends.
Experiment readouts go wrong in predictable ways: analysing a test whose traffic split is broken (a sample ratio mismatch usually means a bug in assignment or logging, and invalidates the result), reporting a p-value without the size and uncertainty of the effect, calling a win after peeking or after testing many metrics and segments, and ignoring guardrails. A good readout checks validity first, then estimates the effect with an interval, then decides against criteria that were set before the test.
Analyse this experiment. Primary metric: .
Only if [GUARDRAILS] is given: Guardrail metrics and thresholds:
- Data quality: run a sample-ratio-mismatch check with a chi-square goodness-of-fit test against the intended split (assume an equal split if none is given, and say so). Treat p < 0.001 as a mismatch. Also note anything else suspicious: very short duration, less than one full weekly cycle, or a metric that is implausibly different.
- If there is a mismatch, stop the effect analysis, give the decision "Do not trust: investigate assignment", and list likely causes to check.
- Primary metric: compute each variant's value, the absolute difference and relative lift, a two-sided 95% confidence interval for the difference (two-proportion z-interval for rates; Welch's t-interval for means), and the p-value. Compare the interval with the minimum detectable or practically meaningful effect if one was given.
- Check the unit of analysis. If the metric's denominator is not the randomisation unit (for example conversion per session or revenue per order while users were randomised), observations are not independent and the naive interval is too narrow. Use per-unit aggregates with the delta method, or ask for per-user data, and say which you did.
- Guardrails: for each, compute the difference and its interval and say whether the interval rules out a breach of the threshold (non-inferiority), shows a breach, or is inconclusive.
- Caveats: multiple variants or metrics (apply a correction such as Holm and say so), early stopping or peeking, novelty effects, segment results (exploratory only), and whether the test was powered for the observed effect.
- Decide, using the first rule that applies:
- Do not trust: the SRM check failed or another data-quality problem invalidates the comparison.
- Stop: the primary metric is worse, or its whole interval lies below the smallest effect worth having (flat, or too small to matter).
- Ship: the interval's lower bound is above zero, the effect is large enough to matter (judged against the stated minimum effect, or say that none was given), and every guardrail passes.
- Iterate: anything else, such as an interval that includes zero but leaves a worthwhile effect possible, or a primary win with a guardrail that is breached or inconclusive.
- Show the formulas and the arithmetic so the reader can check them. If you can run code, compute the numbers with it and say so; otherwise compute carefully by hand and round only in the final line.
- Use only the numbers provided. If you need a value that is missing (for example standard deviations for a mean metric, or the number of users per variant), ask for it and do not estimate it. In that case write "Cannot decide yet" under Decision, name the missing values, and complete only the sections the given numbers support.
- Never call a result significant or not on the p-value alone; always report the interval.
- Treat segment results and secondary metrics as hypotheses for a follow-up test, not as grounds to ship.
- Use the decision words exactly: Ship, Iterate, Stop, Do not trust, or Cannot decide yet.
Decision
The decision word, then two or three sentences on why.
Data quality
The SRM result (observed vs expected counts, chi-square, p) and any other warnings.
Primary metric
A table: variant | n | value | absolute difference | relative lift | 95% CI | p-value. After a failed SRM check, write "Not analysed: sample ratio mismatch" here and under Guardrails.
Guardrails
A table: metric | difference | 95% CI | threshold | status (pass / breach / inconclusive).
Caveats
Bullets.
Calculations
The formulas and arithmetic.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Data analysis
- category
- Statistics
- level
- Intermediate
- made for
- Data analyst, Data scientist, Product manager, Marketer
- risk
- read-only
- version
- v1.1.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install analyze-ab-test-results --target claude-codenpx skills add hermes-hq/hodios-dist --skill analyze-ab-test-results -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-data-analysis@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of StatisticsConsulting statistician
Consulting statistician who asks how the data were produced before analysing them, chooses methods that fit the question, checks assumptions and refuses to over-claim. Use for any data analysis.
statisticianDesign an A/B test
Designs an A/B test plan with a hypothesis, primary and guardrail metrics, minimum detectable effect, sample size, duration, randomisation unit, stop rules and an analysis plan.
design-ab-testBuild a KPI driver tree
Decomposes a top-line metric into a driver tree with exact formulas, definitions and owners, so a change in the metric can be traced to the input that moved. Use for metric design and reviews.
build-kpi-treeCalculate sample size
Computes the sample size or statistical power for an experiment or survey, shows the formula and assumptions, and gives a sensitivity table. Use before launching an A/B test, study or survey.
calculate-sample-sizeCheck an analysis for pitfalls
Reviews an analysis for statistical pitfalls such as Simpson's paradox, p-hacking, survivorship, base rates and causal over-claims before it is shared. Use as a pre-publication review.
check-analysis-for-pitfallsChoose a statistical test
Picks the right statistical test for a research question and data shape, explains its assumptions and how to check them, and gives code to run it. Use before testing a difference or relationship.
choose-statistical-test