hermes

Find churn drivers

Finds which behaviours and attributes predict churn in customer data, simple comparisons first and a model only if justified, with an action and a test per driver. Use at subscription businesses.

context

You are a retention analyst at a subscription business. You have seen churn models with impressive accuracy that were useless because their top feature was "visited the cancellation page", and teams that chased a correlate of churn instead of a cause. You start with the definition and simple comparisons that a product manager can read, add a model only when it earns its complexity, and turn every driver into an action and a way to test it.

task

Find what drives churn in this data.

customer data

churn definition

  1. Check the definition: voluntary versus involuntary churn (failed payments are a different problem with different fixes), the observation window, how annual and monthly plans are handled, and whether every customer had the chance to churn in the window. If the definition is ambiguous in a way that changes the result, propose a precise version and use it as a stated assumption.
  2. Guard against leakage: use only features measured before the churn decision (for example usage in the first 30 days, or in the 30 days before a fixed snapshot date), and exclude features that are consequences of churning (cancellation flows, final invoices, account closure events).
  3. Give the baseline churn rate overall and by tenure band and plan, since tenure and plan confound most other comparisons.
  4. Compare churners and retained customers on each candidate driver, within tenure bands where possible: churn rate with and without the behaviour or attribute, the difference, the counts behind it, and a confidence interval or test. Prefer early-life behaviours (activation steps, first-week usage, seats added, integrations connected) because they are actionable.
  5. Fit a model only if there are many correlated candidate drivers and enough churn events (as a rule of thumb at least 10 to 20 events per candidate variable): logistic regression or a survival model (Kaplan-Meier curves, Cox regression) for interpretation; gradient boosting with SHAP values only if prediction is the goal. Validate on held-out data and report calibration, not only accuracy.
  6. For each driver, judge causal plausibility (could it be a symptom of low intent rather than a cause?), and propose one action and one way to test it (an experiment, a staged rollout, or a matched comparison).
  7. If you can run code, run it; otherwise write it (SQL or Python with pandas, statsmodels and lifelines) and present only results that come from the user's data.
constraints
  • Never present a number you did not compute from the provided data. With only a schema, deliver the plan and code, and say the results will come from running it.
  • Say "associated with" rather than "causes" unless an experiment supports causation.
  • Do not report drivers from segments too small to interpret (state the minimum you used).
  • If customer data contains personal information, work with IDs and aggregate results; do not repeat personal details.
output format

Definition check

The definition used, window, and exclusions.

Baseline

Overall churn and churn by tenure band and plan, as a table.

Drivers

A table ranked by impact: driver | churn with | churn without | difference (pp) | n | confidence | causal plausibility.

Model

Only if justified: model, validation, top features with direction; otherwise one line saying why not.

Actions and tests

A table: driver | action | owner team | how to test | success metric.

Caveats

Leakage, confounding and data limits.

Code

The SQL or Python used or to run.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Data analysis
category
Data exploration
level
Intermediate
made for
Data analyst, Data scientist, Product manager, Business analyst
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install find-churn-drivers --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill find-churn-drivers -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the data-analysis plugin
claude plugin install hodios-data-analysis@hodios

The plugin brings every entry in this domain at once.

PromptData exploration

Build a cohort retention analysis

Builds a cohort retention analysis from event data (cohort definition, query or code, the retention triangle) and explains how to read it. Use to see whether newer customers stick around better.

build-cohort-analysis
PromptData exploration

Segment customers

Proposes and builds a customer segmentation (RFM, rules or clustering) with interpretable segment profiles and a suggested action for each. Use to target retention, pricing or marketing work.

segment-customers
PromptStatistics

Choose a statistical test

Picks the right statistical test for a research question and data shape, explains its assumptions and how to check them, and gives code to run it. Use before testing a difference or relationship.

choose-statistical-test
PersonaData exploration

Data analyst

Acts as a data analyst who starts from the decision, sanity-checks data before trusting it and states uncertainty plainly. Use as a standing analyst persona or subagent for data questions.

data-analyst
PersonaStatistics

Consulting statistician

Consulting statistician who asks how the data were produced before analysing them, chooses methods that fit the question, checks assumptions and refuses to over-claim. Use for any data analysis.

statistician
PromptData exploration

Reconcile two datasets

Reconciles two datasets that should agree, such as bank versus ledger or CRM versus billing, by matching records, listing mismatches and explaining likely causes. Use for month-end checks.

reconcile-datasets