hermes

Analyse A/B test results

Analyses A/B test results with a sample-ratio-mismatch check, effect sizes, confidence intervals and guardrail metrics, ending in a ship, iterate or stop call. Use when an experiment ends.

context

Experiment readouts go wrong in predictable ways: analysing a test whose traffic split is broken (a sample ratio mismatch usually means a bug in assignment or logging, and invalidates the result), reporting a p-value without the size and uncertainty of the effect, calling a win after peeking or after testing many metrics and segments, and ignoring guardrails. A good readout checks validity first, then estimates the effect with an interval, then decides against criteria that were set before the test.

task

Analyse this experiment. Primary metric: .

results

Only if [GUARDRAILS] is given: Guardrail metrics and thresholds:

  1. Data quality: run a sample-ratio-mismatch check with a chi-square goodness-of-fit test against the intended split (assume an equal split if none is given, and say so). Treat p < 0.001 as a mismatch. Also note anything else suspicious: very short duration, less than one full weekly cycle, or a metric that is implausibly different.
  2. If there is a mismatch, stop the effect analysis, give the decision "Do not trust: investigate assignment", and list likely causes to check.
  3. Primary metric: compute each variant's value, the absolute difference and relative lift, a two-sided 95% confidence interval for the difference (two-proportion z-interval for rates; Welch's t-interval for means), and the p-value. Compare the interval with the minimum detectable or practically meaningful effect if one was given.
  • Check the unit of analysis. If the metric's denominator is not the randomisation unit (for example conversion per session or revenue per order while users were randomised), observations are not independent and the naive interval is too narrow. Use per-unit aggregates with the delta method, or ask for per-user data, and say which you did.
  1. Guardrails: for each, compute the difference and its interval and say whether the interval rules out a breach of the threshold (non-inferiority), shows a breach, or is inconclusive.
  2. Caveats: multiple variants or metrics (apply a correction such as Holm and say so), early stopping or peeking, novelty effects, segment results (exploratory only), and whether the test was powered for the observed effect.
  3. Decide, using the first rule that applies:
  • Do not trust: the SRM check failed or another data-quality problem invalidates the comparison.
  • Stop: the primary metric is worse, or its whole interval lies below the smallest effect worth having (flat, or too small to matter).
  • Ship: the interval's lower bound is above zero, the effect is large enough to matter (judged against the stated minimum effect, or say that none was given), and every guardrail passes.
  • Iterate: anything else, such as an interval that includes zero but leaves a worthwhile effect possible, or a primary win with a guardrail that is breached or inconclusive.
constraints
  • Show the formulas and the arithmetic so the reader can check them. If you can run code, compute the numbers with it and say so; otherwise compute carefully by hand and round only in the final line.
  • Use only the numbers provided. If you need a value that is missing (for example standard deviations for a mean metric, or the number of users per variant), ask for it and do not estimate it. In that case write "Cannot decide yet" under Decision, name the missing values, and complete only the sections the given numbers support.
  • Never call a result significant or not on the p-value alone; always report the interval.
  • Treat segment results and secondary metrics as hypotheses for a follow-up test, not as grounds to ship.
  • Use the decision words exactly: Ship, Iterate, Stop, Do not trust, or Cannot decide yet.
output format

Decision

The decision word, then two or three sentences on why.

Data quality

The SRM result (observed vs expected counts, chi-square, p) and any other warnings.

Primary metric

A table: variant | n | value | absolute difference | relative lift | 95% CI | p-value. After a failed SRM check, write "Not analysed: sample ratio mismatch" here and under Guardrails.

Guardrails

A table: metric | difference | 95% CI | threshold | status (pass / breach / inconclusive).

Caveats

Bullets.

Calculations

The formulas and arithmetic.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Data analysis
category
Statistics
level
Intermediate
made for
Data analyst, Data scientist, Product manager, Marketer
risk
read-only
version
v1.1.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install analyze-ab-test-results --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill analyze-ab-test-results -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the data-analysis plugin
claude plugin install hodios-data-analysis@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Statistics
PersonaStatistics

Consulting statistician

Consulting statistician who asks how the data were produced before analysing them, chooses methods that fit the question, checks assumptions and refuses to over-claim. Use for any data analysis.

statistician
PromptProduct metrics

Design an A/B test

Designs an A/B test plan with a hypothesis, primary and guardrail metrics, minimum detectable effect, sample size, duration, randomisation unit, stop rules and an analysis plan.

design-ab-test
PromptReporting

Build a KPI driver tree

Decomposes a top-line metric into a driver tree with exact formulas, definitions and owners, so a change in the metric can be traced to the input that moved. Use for metric design and reviews.

build-kpi-tree
PromptStatistics

Calculate sample size

Computes the sample size or statistical power for an experiment or survey, shows the formula and assumptions, and gives a sensitivity table. Use before launching an A/B test, study or survey.

calculate-sample-size
PromptStatistics

Check an analysis for pitfalls

Reviews an analysis for statistical pitfalls such as Simpson's paradox, p-hacking, survivorship, base rates and causal over-claims before it is shared. Use as a pre-publication review.

check-analysis-for-pitfalls
PromptStatistics

Choose a statistical test

Picks the right statistical test for a research question and data shape, explains its assumptions and how to check them, and gives code to run it. Use before testing a difference or relationship.

choose-statistical-test