Write property-based tests
Finds the invariants a function must keep and writes property-based tests with generators that shrink well. Use when example-based tests miss edge cases in parsers, encoders or pure logic.
Property-based tests state a rule that must hold for every valid input and let a generator search for a counterexample, then shrink it to the smallest failing case. They find the bugs example tests miss, but only when the property is genuinely true of the specification (not a restatement of the implementation) and the generators produce valid, varied, shrinkable inputs. A property that re-implements the function proves nothing; a generator that filters away 90% of its draws is slow and shrinks badly.
Write property-based tests for:
Library: Only if [FRAMEWORK] is given: . If no library is named, detect it from the project's manifests and existing tests (Hypothesis for Python, fast-check for JavaScript and TypeScript, proptest for Rust, jqwik for Java, FsCheck for .NET, rapid for Go; the standard library's testing/quick is frozen and shrinks nothing). If none is installed, pick the standard one for the language and say how to add it.
- Read the code and state its contract: valid inputs, outputs, errors it may raise, and side effects. If the contract is ambiguous (for example, what happens on empty input), ask or state the assumption you test against.
- Find candidate properties, preferring these patterns:
- round-trip: decode(encode(x)) == x, parse(print(x)) == x;
- invariants: output is sorted, length preserved, total conserved, no duplicates, within bounds;
- idempotence: f(f(x)) == f(x);
- oracle or model: agrees with a simpler, obviously correct implementation or an in-memory model of a stateful system;
- metamorphic: a known change to the input causes a predictable change to the output;
- algebraic: commutativity, associativity, identity elements where the domain promises them;
- robustness: never crashes or hangs on any input of the right type, and fails only with documented errors. Keep only properties that follow from the contract. Discard any that just mirror the implementation.
- Build generators from the domain, not from raw types: construct valid values directly (map, compose, build strategies) instead of generating anything and filtering. Include the edge values the type allows: empty, single element, zero, negative, maximum sizes, Unicode beyond ASCII, NaN and infinities for floats where relevant. Bound sizes so a run stays fast.
- Write the tests in the project's style and test runner. Make failures reproducible: rely on the library's seed reporting and example database or replay, and add any shrunk counterexample you discover as an explicit regression example.
- If you can run the tests, do so and report the result. If a property fails, report the minimal counterexample and whether the bug is in the code or in your property. Do not change the code under test.
- Every property must name the contract clause it checks. No property may call the function under test to compute its own expected value.
- Avoid filter or assume calls that reject more than a small fraction of draws; restructure the generator instead.
- Keep default example counts unless there is a reason to change them, and say why if you do.
- Do not fix bugs you find; report them.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
Properties
Table: Property | Pattern | Contract clause it checks.
Generators
One line per generator: what it builds and which edge values it covers.
Tests
The complete test file in one code block, with imports.
Counterexamples
Shrunk failing inputs with a one-line diagnosis each, or "None found" with the number of examples run. If you could not run the tests, say so.
How to run
The exact command, including how to replay a failure from its seed.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Testing
- level
- Intermediate
- made for
- Software engineer, QA / test engineer
- needs
- repo-read
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install write-property-based-tests --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-property-based-tests -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of TestingTest engineer
Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.
test-engineerReview test quality
Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.
review-test-qualityTest-writing rules
Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.
test-writing-rulesWrite a test plan
Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.
write-test-planAdd characterization tests to legacy code
Pins down what untested legacy code does today with characterization and golden-master tests, bugs included, so it can be changed safely. Use before refactoring or modifying code with no tests.
add-characterization-testsAdd a regression test for a bug
Writes the smallest test that fails on the buggy code and passes with the fix, and proves both by running it. Use after fixing a bug, or before fixing one, so it cannot return.
add-regression-test