# Hodios paste pack: Product metrics

Everything in Product metrics from Hodios, the open prompt library by Hermes IDE: 10 entries, catalog 2026.1003.0.

Every entry is dedicated to the public domain under CC0 1.0. Copy, change and share them freely, no attribution needed.

Browse and search the library at https://hermes-ide.com/prompts

## How to use

Find an entry below and copy the text inside its block into ChatGPT, claude.ai or any chat. Replace each [PLACEHOLDER] with your own material. Personas, rules and styles work best as custom instructions or project instructions.

## Contents

- Product metrics
  - [Analyse a conversion funnel](#analyze-conversion-funnel) (prompt)
  - [Build a growth experiment backlog](#build-experiment-backlog) (prompt)
  - [Define a north star metric](#define-north-star-metric) (prompt)
  - [Define an activation metric](#define-activation-metric) (prompt)
  - [Define feature success metrics](#define-feature-success-metrics) (prompt)
  - [Design an A/B test](#design-ab-test) (prompt)
  - [Diagnose a metric drop](#diagnose-metric-drop) (prompt)
  - [Estimate a feature's impact](#estimate-feature-impact) (prompt)
  - [Review launch results](#review-launch-results) (prompt)
  - [Write an analytics tracking plan](#write-tracking-plan) (prompt)

---

<a id="analyze-conversion-funnel"></a>

## Analyse a conversion funnel

`analyze-conversion-funnel` · prompt · Product metrics · https://hermes-ide.com/prompts/analyze-conversion-funnel

Analyses a conversion funnel step by step to find the biggest leak, the segments where it differs, likely causes and the experiments or fixes worth trying first. For PMs and growth teams.

````markdown
<context>
You are a product analyst who works with growth teams. Funnel analysis goes wrong when step counts are compared without checking definitions (users versus sessions, strict versus loose ordering, different time windows), when the "biggest drop" is judged by percentage alone while a later step loses more users who matter more, when an average hides one segment that is broken, and when causes are asserted without evidence. Your job is to find where the funnel leaks most in a way the team can act on, show the arithmetic, and propose the fixes and tests worth trying first.
</context>

<task>
Funnel data:

<funnel_data>
[FUNNEL_DATA]
</funnel_data>

1. Check the data before analysing: the unit (users, sessions, accounts), whether steps are strictly ordered, the conversion window, the date range, whether any step count is higher than the previous one (a sign of loose ordering or tracking issues), and recent tracking or product changes. List anything that makes the numbers unreliable, and keep going only with what can be trusted.
2. Compute, for each step: the count, conversion from the previous step, conversion from the top, and the number of users lost. Show the arithmetic.
3. Find the biggest leak, judged on three things together: users lost at the step, how far that step's conversion is from what the team can plausibly reach (from comparable segments, past periods or the input, not from invented industry benchmarks), and the value of the users lost (later steps usually lose more qualified users). Explain the choice.
4. If segment data is present, compare conversion at the leaky step (and overall) across segments. Highlight segments that differ meaningfully, with their sample sizes; ignore differences that small samples could explain and say so. Look for mix shift: an overall change caused by more traffic from a weaker segment rather than a change in behaviour.
5. List likely causes for the leak, grouped as: tracking or data artefact, technical problem (errors, speed, a specific browser or device), usability friction, intent or expectation mismatch (traffic that was never going to convert, a promise the page does not keep), and pricing or trust. For each cause, give the evidence for and against from the data and flow, and how to check it quickly.
6. Propose what to do next: quick fixes for obvious defects, and two to four experiments, each with a hypothesis, the change, the primary metric, a rough expected effect (stated as an assumption) and how to test it. Order by expected impact relative to effort.
7. List the data to pull next to confirm or rule out the top causes.
</task>

<constraints>
- Show every calculation; round percentages to one decimal place.
- Do not invent benchmarks, segment data or causes presented as facts. Label hypotheses as hypotheses.
- Flag small samples (for example fewer than about 100 users at a step in a segment) as directional.
- If only two steps are given, say the analysis is limited and suggest the intermediate steps to instrument.
</constraints>

<output_format>
## Data check
Bullets, ending with what is trusted.

## Funnel
Table: step | count | step conversion | conversion from top | users lost.

## Biggest leak
The step and the reasoning in three to five sentences.

## Segments
Table: segment | n at step | conversion at leaky step | overall conversion | note. Or "No segment data provided".

## Likely causes
Table: cause | category | evidence for | evidence against | quick check.

## What to do next
Quick fixes, then experiments: hypothesis | change | metric | expected effect (assumed) | effort.

## Data to pull next
Bullets.
</output_format>
````

---

<a id="build-experiment-backlog"></a>

## Build a growth experiment backlog

`build-experiment-backlog` · prompt · Product metrics · https://hermes-ide.com/prompts/build-experiment-backlog

Builds a ranked growth experiment backlog from a funnel and ideas, with hypothesis, metric, effort, expected impact, minimum sample and run time per test, and flags untestable ideas.

````markdown
<context>
You are a growth lead who runs an experimentation programme. Backlogs go wrong in three ways: they rank by excitement instead of by impact on the weakest step, they include tests that cannot reach significance with the traffic available, and their hypotheses are restated ideas ("Make the button green") with no reason or metric. You rank by expected value and testability, and you do the sample-size arithmetic before anyone builds a variant.

Sample size rule of thumb for a two-variant test on a conversion rate, at 5% two-sided significance and 80% power (Lehr's rule): n per variant ≈ 16 × p × (1 − p) / d², where p is the baseline rate and d is the absolute lift you want to detect (minimum detectable effect). Run time = (n × number of variants) / weekly eligible traffic, rounded up to whole weeks, and never under one full week (two is better) so weekday effects even out.
</context>

<task>
<funnel_data>
[FUNNEL_DATA]
</funnel_data>

If the funnel has no counts or rates at all, ask for them and stop.

1. **Funnel diagnosis.** Compute step-to-step conversion and the absolute drop-off at each step. Name the two or three steps where a realistic improvement would add the most completed conversions at the end of the funnel, and why.
2. **Ideas.** Use the team's ideas. If fewer than about eight, or none target the weakest steps, add proposals and label them "proposed". Merge duplicates.
3. **Score each idea:**
   - Hypothesis: "Because we observed [evidence], we believe [change] for [users] will raise [metric], because [mechanism]." Evidence that is an assumption is labelled as such.
   - Primary metric and the funnel step it moves; one guardrail.
   - Expected impact: a relative lift range (for example 3-8%) with the reasoning, and the extra end-of-funnel conversions per month at the midpoint.
   - Confidence: high, medium or low, based on the evidence.
   - Effort: S, M or L (days of design and engineering, as a stated assumption).
   - Minimum sample per variant and run time, using the rule above with the baseline for that step and the midpoint lift converted to an absolute d. Show the numbers.
4. **Rank.** Score = expected extra conversions per month × confidence weight (high 1, medium 0.6, low 0.3) ÷ effort weight (S 1, M 2, L 4), and order by score. Any test that needs more than eight weeks to run leaves the ranked backlog and goes to step 6.
5. **Top test cards.** For the top three, a card: hypothesis, variants, audience and allocation, primary metric, guardrails, sample and duration, the decision rule, and what to do with each outcome.
6. **Not testable as an A/B test.** Ideas that cannot reach the needed sample within eight weeks: say why and what to do instead (make a bolder change with a larger expected lift, test on a higher-traffic step, use a before-and-after with a holdout, qualitative tests, or just ship it if it is low risk and clearly better).
</task>

<constraints>
- Every computed number shows its inputs. Do not invent baselines or traffic: if a step's traffic is missing, write the formula and mark the run time "needs traffic".
- Expected lifts are estimates; keep them modest (most tests win small or not at all) and never present them as forecasts.
- One primary metric per test. No test changes several unrelated things at once unless it is labelled a bundle test.
- No dark patterns in proposed ideas: no fake urgency, hidden costs or pre-ticked consent.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Funnel diagnosis
| Step | Users | Step conversion | Drop-off |
Then two or three bullets on where to focus.

## Ranked backlog
| Rank | Idea | Step and metric | Hypothesis (short) | Expected lift | Extra conversions/month | Confidence | Effort | Score | n per variant | Run time |

## Top test cards
One card per test as a short bulleted block.

## Not testable as an A/B test
## Assumptions
</output_format>

<examples>
<example>
Baseline checkout completion p = 0.40, target relative lift 5% → d = 0.02. n ≈ 16 × 0.40 × 0.60 / 0.0004 = 9,600 per variant. With 6,000 eligible users a week and two variants: 19,200 / 6,000 = 3.2 → 4 weeks.
</example>
</examples>
````

---

<a id="define-north-star-metric"></a>

## Define a north star metric

`define-north-star-metric` · prompt · Product metrics · https://hermes-ide.com/prompts/define-north-star-metric

Proposes a north star metric with input metrics and guardrails, tests it against the value users actually get, and shows the rejected candidates. Use when setting product goals.

````markdown
<context>
You are a product analytics leader who helps teams choose a north star metric. A good north star captures the value customers get from the product, leads revenue rather than being revenue, can be influenced by the team, is understandable by everyone, and moves within weeks rather than years. It is decomposed into a few input metrics that teams can own. It fails when it is a vanity count (signups, page views), a lagging financial number, or something that can rise while users are worse off, such as time spent on a product meant to save time.
</context>

<task>
Product:

<product>
[PRODUCT]
</product>

Business model:

<business_model>
[BUSINESS_MODEL]
</business_model>

1. Identify the core value exchange: what the user gets, the action that delivers it, and the natural frequency of that action (daily, weekly, monthly, a few times a year). Classify the product's game: attention (time and engagement are the value), transaction (completed exchanges are the value) or productivity (work done efficiently is the value).
2. Propose three or four candidate north stars. Score each against: reflects customer value, leading indicator of revenue, actionable by teams, understandable, measurable now, and resistant to gaming. Prefer metrics that count users or units achieving value in a period (for example "weekly teams that complete at least 3 shared projects") over raw totals.
3. Recommend one. Define it precisely: the unit, the qualifying action and threshold, the time window, and what is excluded (internal users, bots, test accounts).
4. Break it into three to five input metrics, using breadth (how many users), depth (how much value per user), frequency (how often) and efficiency (how quickly or easily). Name which team could own each.
5. Add guardrail metrics that catch harmful ways to move the north star (for example support contacts, refunds, unsubscribes, quality ratings, margin).
6. Run the value check: describe at least two ways the metric could go up while customers are worse off or the business is weaker, and show how the guardrails or the definition prevent it.
7. Explain how to roll it out: data needed, a baseline to establish, review cadence, and when to revisit the choice.
</task>

<constraints>
- Do not choose revenue, signups, downloads or page views as the north star; they may appear as guardrails or business outcomes.
- The metric's time window must match the natural frequency of use; a monthly-use product must not have a daily active metric.
- If the product description is too thin to identify the core value, ask up to three questions and stop.
- Do not invent current values or benchmarks; say what must be measured.
</constraints>

<output_format>
## Recommendation
The north star in one line, then its precise definition as bullets (unit, qualifying action, window, exclusions).

## Candidates considered
Table: candidate | value | leads revenue | actionable | understandable | measurable | gaming risk | verdict.

## Metric tree
An indented tree: north star, then input metrics with their type (breadth, depth, frequency, efficiency) and owning team.

## Guardrails
Table: guardrail | what harm it catches | alert threshold to set.

## Value check
Bullets: the failure mode and the protection.

## How to roll it out
Up to five bullets.
</output_format>
````

---

<a id="define-activation-metric"></a>

## Define an activation metric

`define-activation-metric` · prompt · Product metrics · https://hermes-ide.com/prompts/define-activation-metric

Finds a product's activation moment from usage and retention data, defines an activation metric with an action, threshold and time window, and plans how to validate it.

````markdown
<context>
You are a product analyst who has defined activation metrics for consumer and B2B products. An activation metric names the early behaviour that separates new users who go on to retain from those who do not, in a form the team can move: "created 3 projects and invited 1 teammate within 7 days of sign-up". It is a leading indicator for onboarding work. Teams get it wrong by picking the action with the highest raw retention lift while only 2% of users do it, by picking something nearly everyone does, by choosing a window so long that it cannot steer onboarding, and by treating a correlation as proof that pushing users to the action will cause retention.
</context>

<task>
<usage_data_summary>
[USAGE_DATA_SUMMARY]
</usage_data_summary>

If the data has no retention or conversion outcome, or no split between users who did and did not do the candidate actions, do not guess: explain what is missing and give the analysis to run (step 6) instead of a recommendation.

1. **Retention outcome.** State the outcome the activation metric predicts (for example "active in week 4", "converted to paid by day 30", "account still active in month 3") and check it fits the product's natural usage frequency. If the data uses a different outcome, use it and note the mismatch.
2. **Candidate actions.** For each candidate action and threshold in the data, compute or extract:
   - Reach: share of new users who reach it in the window.
   - Retention if reached and if not reached, and the lift between them.
   - Coverage: share of retained users who reached it (how much of retention it explains).
   - Precision: share of users who reached it who retained.
   Show the calculation when you derive a number. Where several thresholds exist (1, 3, 5 projects), find where the retention gain flattens.
3. **Recommended activation metric.** Pick the action, threshold and window that best balance precision and coverage while being reachable early enough to steer onboarding. Prefer an action that reflects receiving value (completing a report, a teammate responding) over setup busywork (filling in a profile). Explain why it beats the runner-up. If two actions together beat either alone, consider a combined definition, but keep it explainable in one sentence.
4. **Metric definition.** A precise spec: name, plain-language definition, numerator, denominator (which sign-up cohort, which exclusions such as test accounts, internal users or invited users), window measured from what event, the events and properties needed, refresh cadence, and an owner placeholder.
5. **Validation plan.** How to check the metric is useful, not just correlated: hold the definition fixed on a later cohort; check it holds across the main segments and acquisition channels; and run at least one onboarding experiment that raises the activation rate, then check whether retention in that test group rises too. Set the result that would make you revise the definition.
6. **Analysis to run.** If the data was insufficient, or to confirm the recommendation, describe the query: cohort, events, windows, outputs per threshold. Use plain pseudo-SQL or step-by-step logic.
7. **Caveats.** Selection effects (motivated users do everything), small samples, seasonality, and how the definition could be gamed.
</task>

<constraints>
- Every number you report comes from the data given or is computed from it with the working shown. Never invent rates or sample sizes.
- Flag any candidate with fewer than about 100 users in either group as too small to rank confidently.
- Describe relationships as associations; causal language is allowed only for experimental results.
- Keep the metric to one sentence a new team member would understand.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Retention outcome
One or two sentences.

## Candidate actions
| Action and threshold | Window | Reach | Retention if reached | Retention if not | Lift | Coverage | Precision | Notes |

## Recommended activation metric
The one-sentence metric in bold, then the reasons and the runner-up.

## Metric definition
Bullets for each spec field.

## Validation plan
Numbered steps with the revise-if condition.

## Caveats
Bullets. Add "## Analysis to run" before Caveats when needed.
</output_format>
````

---

<a id="define-feature-success-metrics"></a>

## Define feature success metrics

`define-feature-success-metrics` · prompt · Product metrics · https://hermes-ide.com/prompts/define-feature-success-metrics

Defines success metrics for a feature using HEART and goals-signals-metrics, with baselines, targets, guardrails, decision rules and the event data needed. Use before building or launching.

````markdown
<context>
You are a product analytics lead. You use Google's HEART framework (Happiness, Engagement, Adoption, Retention, Task success) to pick which dimensions of user experience matter for a feature, and the goals-signals-metrics process to turn each one into something measurable: a goal (what success looks like for users), a signal (the behaviour or attitude that shows it), and a metric (the number you track). Teams misuse both by filling in all five dimensions with vanity counts, setting targets with no baseline, declaring success on a metric the feature could not move, and forgetting what the feature might break.
</context>

<task>
<feature>
[FEATURE]
</feature>

If the feature description does not say what user problem it solves or who it is for, ask and stop.

1. **Goals.** Write two or three user-centred goals and the business goal they serve. If goals were not given, propose them and mark them "to confirm".
2. **Metrics.** Choose the two to four HEART dimensions that matter for this feature and explain why the others are left out. For each chosen dimension, give goal, signal, metric (exact formula with numerator, denominator and time window), baseline (from the input, or "unknown: measure for N weeks before launch"), target with time frame and the reasoning behind it, and data source.
3. **Primary metric and decision rule.** Pick one metric the launch decision rests on, and write the rule: "Ship to everyone if X rises by at least Y within Z, with no guardrail breached; iterate if…; roll back if…". Prefer a metric the feature directly moves over a lagging company metric.
4. **Guardrails.** Two to four metrics that must not get worse (for example support contacts, latency, conversion of a nearby flow, unsubscribes, revenue per user), each with its tolerance.
5. **Event data needed.** The events and properties to instrument, with when each fires and which metric uses it. Note any event that already exists according to the input.
6. **Readout plan.** How the effect will be measured (A/B test, staged rollout with holdout, or before-and-after with its weaknesses stated), when to read it (early health check, then the decision date), and who decides.
7. **Open questions.** What must be confirmed before launch.
</task>

<constraints>
- Do not invent baselines. Targets without a baseline are expressed as relative change and flagged for revision once the baseline is known.
- Every metric must be computable from named events or a named data source; drop any that cannot.
- Happiness metrics from surveys need a sample size and a timing (for example in-product survey after the third use); do not rely on them alone for the decision.
- Keep metric names unambiguous: "weekly active users of X" must say what counts as active.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Goals
Bullets.

## Metrics
| HEART dimension | Goal | Signal | Metric (formula, window) | Baseline | Target | Source |
Then one line on the dimensions left out.

## Primary metric and decision rule
## Guardrails
| Metric | Tolerance | Why it could move |
## Event data needed
| Event | Fires when | Properties | Used by |
## Readout plan
## Open questions
</output_format>
````

---

<a id="design-ab-test"></a>

## Design an A/B test

`design-ab-test` · prompt · Product metrics · https://hermes-ide.com/prompts/design-ab-test

Designs an A/B test plan with a hypothesis, primary and guardrail metrics, minimum detectable effect, sample size, duration, randomisation unit, stop rules and an analysis plan.

````markdown
<context>
You are an experimentation lead who reviews test plans before they launch. Most failed A/B tests were decided before they started: a vague hypothesis, a primary metric the change cannot move, too little traffic to detect a realistic effect, the wrong randomisation unit, or a team that peeks daily and stops on the first good day. A good plan is written and agreed before launch, so the result cannot be reinterpreted afterwards.

Primary metric: [PRIMARY_METRIC]


</context>

<task>
Change to test:

<change>
[CHANGE]
</change>

1. Write the hypothesis: "Because [evidence], we believe [change] for [population] will [increase or decrease] [primary metric] by at least [MDE], because [mechanism]."
2. Check the primary metric: it should be sensitive to the change, measured per randomisation unit, and tied to value. If it is far downstream of the change (for example revenue for a button colour), propose a closer metric and keep the original as secondary.
3. Choose two to four guardrail metrics that must not get worse (for example revenue per user, refunds, latency, unsubscribes, support contacts) and any secondary metrics to explain the result.
4. Choose the randomisation unit (user, account, session, device or cluster) and explain why. Use the account or cluster when users interact or share state; note the risk of interference between groups. Define who is eligible and when they are counted (trigger at exposure, not at login, where possible).
5. Set the minimum detectable effect: the smallest change worth shipping. If the user did not give one, propose it with reasoning.
6. Compute the sample size per arm for alpha 0.05 two-sided and 80% power, and show the working. For proportions: n per arm = (1.96 + 0.84)^2 x [p1(1 - p1) + p2(1 - p2)] / (p2 - p1)^2. For means: n per arm = 2 x (1.96 + 0.84)^2 x sd^2 / delta^2. If the baseline is missing, ask for it (and say where to find it) and give the formula ready to fill in. Convert to an enrolment period using the traffic, round up to whole weeks, and set a minimum of one full week. If the metric has a measurement window (for example conversion within 30 days), add that window after the last user enrols to get the time until the result can be read. If the duration is impractical, give the levers: larger MDE, closer metric, variance reduction such as CUPED, more traffic or fewer arms.
7. Write stop rules decided in advance: run to the planned sample unless a guardrail breaches a stated threshold or there is a sample ratio mismatch; no stopping early for a win unless a sequential method is used and named.
8. Write the analysis plan: the test to use, how to handle multiple metrics or arms, the segments you will look at (pre-declared, few), and the decision rule (ship, iterate, or do not ship) for each outcome.
9. List risks and pre-launch checks: tracking verified in both arms, an A/A or SRM check, novelty or learning effects, seasonality and holidays during the window, and other experiments on the same surface.
</task>

<constraints>
- Show every number you use and where it came from (given or assumed). Never invent a baseline rate or variance.
- Keep z-values explicit (1.96 and 0.84) and round sample sizes up.
- Do not recommend peeking-based decisions. If the team needs early reads, recommend a sequential testing method instead.
- If the change touches pricing, consent, or vulnerable users, note any ethical or legal review needed before testing.
</constraints>

<output_format>
## Hypothesis
One sentence in the template above.

## Metrics
Table: metric | role (primary, guardrail, secondary) | definition | direction | threshold.

## Design
Bullets: randomisation unit, eligibility and trigger, arms and split, exclusions.

## Sample size and duration
The MDE, the formula with numbers substituted, n per arm, total, days, and the planned run length in whole weeks.

## Stop rules
Bullets.

## Analysis plan
Bullets, ending with the decision rule.

## Risks and pre-launch checks
A checklist.
</output_format>
````

---

<a id="diagnose-metric-drop"></a>

## Diagnose a metric drop

`diagnose-metric-drop` · prompt · Product metrics · https://hermes-ide.com/prompts/diagnose-metric-drop

Investigates a drop in a product metric with a structured tree (data and tracking, segments, platforms, releases, external factors), ranks the hypotheses and gives the queries to run.

````markdown
<context>
You are a senior product analyst who gets paged when a key metric drops. You have learned that the most common causes are boring: broken tracking, a pipeline delay, a definition change, a mix shift in traffic, or a bad release on one platform. You check whether the drop is real before explaining it, decompose it before theorising, and rank hypotheses by likelihood and cost to check, so the team finds the cause in hours rather than days.

Metric: [METRIC]
</context>

<task>
What changed:

<change>
[CHANGE]
</change>


1. First read: size the drop against normal variation (same weekday last weeks, same period last year), and say whether it is sudden (a step, usually a release, outage or tracking change) or gradual (usually mix, seasonality or product-market change). If key facts are missing (the definition, the comparison period, the size), list them, and continue with what you have.
2. Build the investigation tree, checking in this order:
   - Is it real? Tracking and instrumentation changes, event schema or SDK updates, pipeline delays or partial loads, definition or filter changes, bot filtering, time zone or calendar effects.
   - Decompose: the numerator versus the denominator; each funnel step that feeds the metric; mix shift (segment shares changed) versus rate change (segments' rates changed).
   - Where is it? Platform, app version, OS or browser, country, acquisition channel, new versus returning, plan or customer tier, cohort.
   - Internal causes: releases and feature flags, experiments, pricing or packaging, marketing spend or campaign ends, emails or notifications stopped, outages or latency, support or policy changes.
   - External causes: seasonality and holidays, competitor moves, platform or app store changes, search algorithm updates, payment provider issues, news or regulation.
3. Rank the top hypotheses by likelihood given the evidence and by cost to check, and for each say what you would expect to see if it is true and if it is false.
4. Write the queries to run, in standard SQL with clearly named placeholder tables and columns (for example events(user_id, event_name, event_time, platform, app_version, country)) that the user must map to their schema. Include: the metric by day for a long enough window, the metric split by each key dimension before and after the change date, the funnel steps, and a mix-versus-rate decomposition.
5. Give a decision guide: if a query shows X, the likely cause is Y and the next step is Z.
6. Write a short holding message for stakeholders: what we know, what we are checking, and when the next update will come.
</task>

<constraints>
- Do not name a cause as confirmed; everything is a hypothesis until a query result supports it.
- Do not invent table names as if they were real; mark them as placeholders to adapt.
- If the metric is a ratio, always check the numerator and denominator separately.
- Prefer checks that take minutes (dashboards, release logs, tracking monitors) before deep analysis.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## First read
Three to five bullets.

## Investigation tree
Indented tree with the checks under each branch.

## Ranked hypotheses
Table: rank | hypothesis | why it fits | evidence if true | evidence if false | cost to check.

## Queries to run
Numbered SQL code blocks, each with one line on what it answers.

## Decision guide
Bullets: if this, then that.

## What to tell stakeholders now
A message of under 100 words.
</output_format>
````

---

<a id="estimate-feature-impact"></a>

## Estimate a feature's impact

`estimate-feature-impact` · prompt · Product metrics · https://hermes-ide.com/prompts/estimate-feature-impact

Sizes a feature's expected impact before building it, with explicit reach, adoption, effect and value assumptions, a low-base-high range and the cheapest way to tighten the estimate.

````markdown
<context>
You are a product manager with strong analytical habits who sizes ideas before the team commits to them. Impact estimates go wrong when they apply an optimistic effect to the whole user base instead of the users who will actually see and use the feature, when one point estimate hides huge uncertainty, when cannibalisation and ramp-up are ignored, and when nobody says which assumption the answer depends on. A useful estimate is a simple driver model with every assumption visible, a range rather than a point, and a clear next step to reduce the biggest uncertainty cheaply.
</context>

<task>
Feature:

<feature>
[FEATURE]
</feature>

Baseline metrics:

<baseline_metrics>
[BASELINE_METRICS]
</baseline_metrics>

1. Name the target metric (for example monthly recurring revenue, 30-day retention, support tickets) and write the impact model as a driver chain, typically: reach (users or accounts in the target segment per period) x exposure (share who encounter the feature) x adoption (share of those who use it) x effect (change in the behaviour per adopter) x value (what that change is worth per unit). Adapt the chain to the feature; keep it to five or six drivers.
2. For each driver, give low, base and high values with the source: given in the baseline, derived from it (show how), or assumed (state the reasoning, for example an analogous feature's adoption). Never present an assumed value as data.
3. Compute the impact for low, base and high scenarios, per month and annualised, showing the arithmetic. Note the ramp-up: how long until adoption reaches the steady state, and what that does to first-year impact.
4. Adjust for second-order effects: cannibalisation of existing behaviour or revenue, effects on other metrics (support load, performance), and novelty effects that fade.
5. Sensitivity: which one or two drivers move the result most between low and high? Show the result if only that driver is at its low value.
6. If the build cost is known, compare: payback period at the base case and whether the low case still clears the bar. If unknown, state the break-even cost at the base case.
7. Propose the cheapest ways to tighten the estimate, aimed at the most sensitive drivers: a data pull, a fake door to measure exposure and adoption, a look at an analogous feature's adoption curve, a handful of customer conversations, or a small experiment. Say what each would cost and which driver it narrows.
8. List caveats in one short list.
</task>

<constraints>
- Show all arithmetic; round results to two significant figures to avoid false precision.
- Effects are per adopter, not per user in the base. Never apply the effect to the whole user base unless exposure and adoption are genuinely 100%.
- If the baseline lacks the numbers needed for a driver (for example no segment size), ask for it and use a clearly labelled placeholder range so the model is still useful.
- Do not inflate the high case to make a feature look good; the high case should be plausible, not best imaginable.
</constraints>

<output_format>
## Impact model
The driver chain as a formula.

## Assumptions
Table: driver | low | base | high | source (given, derived, assumed) | reasoning.

## Estimate
Table: scenario | monthly impact | annualised | first-year with ramp-up. Then the arithmetic for the base case.

## Sensitivity
Two or three sentences.

## Is it worth it
Payback or break-even.

## Cheapest ways to tighten the estimate
Table: action | driver narrowed | cost | time.

## Caveats
Bullets.
</output_format>
````

---

<a id="review-launch-results"></a>

## Review launch results

`review-launch-results` · prompt · Product metrics · https://hermes-ide.com/prompts/review-launch-results

Reviews a launched feature against its success criteria, separates real signal from noise and novelty, and recommends whether to iterate, scale or roll back, with the reasoning.

````markdown
<context>
You are a product leader running a post-launch review. Launch reviews go wrong in two directions: teams declare victory on a noisy uptick or a novelty spike, or they quietly move the goalposts to whatever metric happened to rise. You judge the launch against the criteria agreed before it shipped, check whether the evidence is strong enough to support a decision, and make a clear recommendation, even when the honest answer is "not enough data yet".
</context>

<task>
Launch goals and success criteria:

<launch_goals>
[LAUNCH_GOALS]
</launch_goals>

Results:

<results>
[RESULTS]
</results>

1. Restate the pre-agreed success criteria. If there were none, say so, and judge against the most reasonable criteria implied by the goals, labelled as reconstructed after the fact.
2. Build a scorecard: each criterion, its target, the actual result, and met, missed or unclear.
3. Assess signal versus noise for each result:
   - Comparison: was there a control group or holdout, or is this before-and-after? Before-and-after comparisons are confounded by seasonality, marketing and other releases; name any that overlap.
   - Size and certainty: sample sizes, confidence intervals or significance if given, and whether the change exceeds normal week-to-week variation.
   - Time: is the window long enough to see past novelty or learning effects, and is the trend rising, stable or fading?
   - Adoption: how many eligible users discovered, tried and kept using the feature; low adoption explains weak overall effects.
   - Data quality: tracking changes or gaps around the launch.
4. Look at guardrails and side effects: support load, performance, cannibalisation of other features, complaints.
5. Recommend one of: scale (roll out further or invest more), iterate (keep it and fix specific problems), hold (keep collecting data until a stated date or sample), or roll back. Give the two or three reasons that decide it and what would change your mind.
6. Capture what the team learned for future launches.
</task>

<constraints>
- Do not change the success criteria after seeing the results. If you suggest a better metric for the future, put it under learnings.
- Do not call a difference real without a comparison and some sense of its variability; say "unclear" instead.
- Use only the numbers provided. Compute differences and relative changes and show them; do not invent confidence intervals.
- Credit qualitative feedback for what it is: useful for why, weak for how many.
</constraints>

<output_format>
## Recommendation
Scale, iterate, hold or roll back, with the deciding reasons in two to four sentences.

## Scorecard
Table: criterion | target | actual | status (met, missed, unclear) | note.

## Signal or noise
Bullets per key result covering comparison, size, time, adoption and data quality.

## What we learned
Bullets.

## Next steps
Numbered actions with an owner placeholder and a date or trigger.
</output_format>
````

---

<a id="write-tracking-plan"></a>

## Write an analytics tracking plan

`write-tracking-plan` · prompt · Product metrics · https://hermes-ide.com/prompts/write-tracking-plan

Writes an analytics tracking plan with consistently named events and properties, when each fires, the question it answers, privacy notes and QA steps. Use when instrumenting a feature.

````markdown
<context>
You are a product analyst who writes tracking plans that engineers can implement and analysts can trust a year later. Tracking goes wrong when events are named inconsistently ("signup", "Sign Up Completed", "user_registered"), when the moment an event fires is ambiguous (button click or successful save?), when critical events are tracked only in the browser where ad blockers and retries distort them, when personal data leaks into properties, and when events are added with no question behind them. A good plan starts from the questions, defines the minimum set of events and properties that answers them, and says exactly how to verify the data before launch.
</context>

<task>
Feature:

<feature>
[FEATURE]
</feature>

Questions to answer:

<questions>
[QUESTIONS]
</questions>

1. Map each question to the metric that answers it (with numerator, denominator and time window) and to the events and properties needed. If a question cannot be answered with event data (for example "why do users leave?"), say so and suggest the right method instead (survey, interviews, session research).
2. Set naming conventions unless existing ones are given: events as Object + Action in past tense ("Invoice Sent", or invoice_sent in snake case), properties in snake_case, consistent IDs (user_id, account_id), and enumerated values listed explicitly. If existing events are listed, reuse and extend them rather than creating near-duplicates.
3. Define the events. For each: name; the exact trigger (which user action or system outcome, and at what moment: on click, on successful server response, on page view); where it is sent from (client or server - prefer server-side for anything involving money, account state or completion of a critical step); properties with type, example value, allowed values and whether required; and the question it serves. Track outcomes (succeeded or failed with a reason), not only attempts.
4. Define user and account (group) properties that segmentation needs, such as plan, signup date, role, company size band, and when they are set or updated.
5. Write metric definitions for the key funnels or rates built from these events, including step order, conversion window and how repeat events are counted.
6. Add privacy notes: no personal data (names, emails, free text, precise location) in event properties unless there is a documented need and consent; respect consent choices before sending; say which properties might be sensitive and how to handle them (hash, bucket or drop).
7. Write the QA plan: test cases per event (action to perform, expected event and properties), checks in a development environment and in the tool's live view, validation of property types and allowed values, comparison of event counts with the source of truth (for example the database), and monitoring after launch for volume drops or schema violations.
8. List open questions for the team.
</task>

<constraints>
- Every event and property must serve a listed question or a stated segmentation need; cut the rest.
- Do not invent the tool's API calls or features; describe the plan in tool-neutral terms and mark anything tool-specific to verify.
- Be exact about trigger moments; "when the user signs up" is not specific enough.
- If the feature description is too thin to define triggers, list what you need (screens, states, success and failure cases) and give a provisional plan.
</constraints>

<output_format>
## Questions to metrics
Table: question | metric (definition) | events and properties needed.

## Naming conventions
Bullets.

## Events
Table: event | trigger (exact moment) | source (client or server) | properties | question served.

Then, per event with properties, a sub-table: property | type | example | allowed values | required.

## User and account properties
Table: property | type | set when | used for.

## Metric definitions
Bullets.

## Privacy
Bullets.

## QA plan
Checklist.

## Open questions
Numbered.
</output_format>
````
