hermes

Add characterization tests to legacy code

Pins down what untested legacy code does today with characterization and golden-master tests, bugs included, so it can be changed safely. Use before refactoring or modifying code with no tests.

context

A characterization test records what the code actually does, not what it should do. It is a safety net for a later change: if a refactor alters any output, a test fails. That means the tests must pin current behaviour exactly, including odd and probably wrong behaviour, and must fail when the behaviour changes. Tests that only check "no exception" or that assert what the author guessed the code does give false confidence.

task

Write characterization tests for: Only if [ENTRY_POINTS] is given: Known entry points:

  1. Find the entry points (from the list above, or from callers in the repository) and test through the highest-level one that is practical to call. Avoid testing private helpers that a refactor will move.
  2. Find the seams that make the code nondeterministic or hard to call: current time, randomness, generated ids, environment, file system, network, database, global state. For each, choose the least invasive way to control it: an existing parameter or injection point first, then a test double at the module boundary, then a minimal seam (extract a parameter with the current value as its default). Name any production change you need; keep it behaviour-preserving.
  3. Choose inputs that exercise every branch you can see: typical values, boundaries, empty and missing values, error paths, and combinations of flags. Read the conditionals to derive them.
  4. Capture current outputs:
  • for small outputs, assert exact values;
  • for large or structured outputs (reports, HTML, JSON, files), write a golden-master or approval test that stores the output in a snapshot file, with scrubbers that normalise timestamps, ids and unordered collections so the snapshot is stable;
  • record side effects too: calls to collaborators, rows written, messages sent, exceptions raised. Derive expected values by running the code where you can. If you cannot run it, derive them by tracing the code and mark those tests "traced, confirm on first run".
  1. Check the net catches change: for each important branch, describe a small mutation (flip a comparison, drop a line) and confirm a test would fail. Add inputs where none would.
constraints
  • Do not fix bugs. Pin the current behaviour and list it under "Suspicious behaviour", with the test name, so a human decides later.
  • Do not refactor production code beyond the minimal seams named in step 2.
  • Name tests by behaviour (returns_zero_discount_when_cart_empty), not by number.
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
  • Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
  • If a test looks wrong, explain why and ask before changing it.
output format

Behaviour inventory

Table: Entry point | Input class | Current output or side effect.

Seams

Bullets: the nondeterminism or dependency, and how the tests control it (including any production change).

Tests

The complete test file or files, with snapshot files if any.

Suspicious behaviour

Table: Behaviour | Test that pins it | Why it looks wrong. Or "None".

Coverage and gaps

Branches covered, branches not covered and why, and the mutations you checked.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Testing
level
Intermediate
made for
Software engineer, Tech lead / staff engineer
needs
repo-read
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install add-characterization-tests --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill add-characterization-tests -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Testing
PromptRefactoring

Plan splitting a large module

Maps the responsibilities and internal dependencies of an oversized file or class and plans its split into cohesive modules, in small steps that keep tests green. Use before breaking up a god class.

split-large-module
PersonaTesting

Test engineer

Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.

test-engineer
PromptTesting

Review test quality

Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.

review-test-quality
RuleTesting

Test-writing rules

Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.

test-writing-rules
PromptTesting

Write a test plan

Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.

write-test-plan
PromptTesting

Add a regression test for a bug

Writes the smallest test that fails on the buggy code and passes with the fix, and proves both by running it. Use after fixing a bug, or before fixing one, so it cannot return.

add-regression-test