Build a cohort retention analysis
Builds a cohort retention analysis from event data (cohort definition, query or code, the retention triangle) and explains how to read it. Use to see whether newer customers stick around better.
You are a product analyst building a cohort retention analysis. A retention triangle answers one question well: are later cohorts behaving better or worse than earlier ones at the same age? It is easy to get wrong in ways that look plausible: counting calendar periods instead of periods since joining, letting the youngest cohorts' incomplete periods look like drops, or mixing a cohort definition with an activity definition that the cohort event itself satisfies.
Build a cohort retention analysis.
Cohort by:
- Define precisely: the cohort event and date for each user (for example first signup), the period length (month or week, matching the cohort grain unless the activity definition says otherwise), period 0, and the retention measure. Decide whether the cohort event itself counts as period-0 activity and say which.
- Decide the retention type and state it: classic or bounded (active in exactly period N) by default; mention unbounded or rolling retention (active in N or later) only if the use case calls for it.
- Write the code. If the data lives in a SQL warehouse, write SQL for the dialect named or implied in the event data (default postgres) using CTEs: cohorts, activity by period, cohort sizes, then the triangle. If it is a file, write pandas. Compute period number as whole periods since the cohort date, not calendar month minus calendar month on raw timestamps without truncation.
- Output the triangle as cohorts in rows, period numbers in columns, values as percentages of cohort size, with the cohort size as its own column.
- Mark cells that are incomplete because the period has not fully elapsed, and exclude them from averages.
- If the event data includes a sample, compute the triangle on the sample to show the shape, labelled as illustrative.
- If the event data lacks a user identifier, a timestamp, or anything that can satisfy the activity definition, say what is missing and stop.
- Never fill missing cohort-period cells with zeros; an unobserved period is not zero retention.
- Users with activity before their cohort date (data errors, imports) are reported as a count, not silently dropped or kept.
- Do not draw conclusions from cohorts smaller than about 30 users without saying the numbers are noisy.
- Keep time zones consistent between the cohort date and activity timestamps; state the assumption.
Definitions
Bullets: cohort, period, period 0, retained, retention type, time zone.
Code
One code block.
Retention triangle
A Markdown table if computed from a sample (labelled illustrative); otherwise the column layout the code produces.
How to read it
Four to six sentences: reading down a column (cohort quality over time), across a row (decay curve), where the curve flattens, and what change would count as meaningful.
Caveats
Bullets specific to this data: incomplete periods, small cohorts, seasonality, definition changes.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Data analysis
- category
- Data exploration
- level
- Intermediate
- made for
- Data analyst, Product manager, Data scientist, Founder / business owner
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install build-cohort-analysis --target claude-codenpx skills add hermes-hq/hodios-dist --skill build-cohort-analysis -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-data-analysis@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Data explorationAnswer a question with SQL
Turns a business question and a schema into an analytical SQL query, states the assumptions behind it and explains how to read the result. Use when you know the question but not the query.
answer-question-with-sqlDefine a metric
Writes a precise metric definition (formula, grain, filters, edge cases, owner, known caveats) so every team computes the number the same way. Use when a metric is disputed or about to be launched.
define-metricReconcile two datasets
Reconciles two datasets that should agree, such as bank versus ledger or CRM versus billing, by matching records, listing mismatches and explaining likely causes. Use for month-end checks.
reconcile-datasetsWrite a dataframe transformation
Writes pandas or polars code for a described transformation with built-in checks on row counts, nulls, key uniqueness and join cardinality. Use when reshaping, joining or aggregating data.
write-dataframe-transformationAnalyse an employee engagement survey
Analyses an employee engagement survey with group scores under minimum-group-size privacy rules, eNPS, comment themes and three priorities to act on. Use after an engagement or pulse survey closes.
analyze-employee-surveyAnalyse location data
Analyses location data for stores, customers or deliveries to find catchments, density and distance patterns, with the method, code and mapping guidance. Use for site, coverage or delivery questions.
analyze-location-data