# Hodios paste pack: Testing

Everything in Testing from Hodios, the open prompt library by Hermes IDE: 13 entries, catalog 2026.1003.0.

Every entry is dedicated to the public domain under CC0 1.0. Copy, change and share them freely, no attribution needed.

Browse and search the library at https://hermes-ide.com/prompts

## How to use

Find an entry below and copy the text inside its block into ChatGPT, claude.ai or any chat. Replace each [PLACEHOLDER] with your own material. Personas, rules and styles work best as custom instructions or project instructions.

## Contents

- Testing
  - [Add a regression test for a bug](#add-regression-test) (prompt)
  - [Add characterization tests to legacy code](#add-characterization-tests) (prompt)
  - [Find and fill the riskiest test gaps](#fill-test-gaps) (prompt)
  - [Fix a flaky test](#fix-flaky-test) (prompt)
  - [Review test quality](#review-test-quality) (prompt)
  - [Test engineer](#test-engineer) (persona)
  - [Test-writing rules](#test-writing-rules) (rule)
  - [Write a resilient end-to-end test](#write-e2e-test) (prompt)
  - [Write a test plan](#write-test-plan) (prompt)
  - [Write consumer-driven contract tests](#write-contract-tests) (prompt)
  - [Write integration tests with real dependencies](#write-integration-tests) (prompt)
  - [Write property-based tests](#write-property-based-tests) (prompt)
  - [Write unit tests](#write-unit-tests) (prompt)

---

<a id="add-regression-test"></a>

## Add a regression test for a bug

`add-regression-test` · prompt · Testing · https://hermes-ide.com/prompts/add-regression-test

Writes the smallest test that fails on the buggy code and passes with the fix, and proves both by running it. Use after fixing a bug, or before fixing one, so it cannot return.

````markdown
<context>
A regression test is only worth its place in the suite if it fails without the fix. Many "regression tests" pass on the broken code too, because they test a neighbouring path or assert too little. The proof is running the test against both versions.
</context>

<task>
Add a regression test for: [BUG]
1. State the bug as one triggering input and one expected result.
2. Find the lowest level where the bug can be observed (unit before integration before end-to-end), and the existing test file where a test for that code belongs.
3. Write one focused test with that input and the expected result. Name it after the behaviour, and reference the issue in a comment if there is one.
4. Prove it:
   - On the code without the fix, the test must fail, and fail for the right reason (the assertion on the bug, not an import or setup error). If the fix is already applied, revert it temporarily, for example with `git stash` or by checking out the parent commit of the fix in a separate worktree.
   - On the code with the fix, the test must pass.
   - If the bug is not fixed yet, the test fails now; report that and leave the fix to the user.
5. Run the surrounding test file or suite to confirm nothing else broke, and restore the work tree to the state you found it in.
</task>

<constraints>
- One bug, one test. Add a second test only for a distinct boundary of the same bug, and say why.
- Do not change production code, except to temporarily revert the fix during the proof.
- Never leave the work tree with the fix reverted or with stashed changes the user did not make.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Test
The file path and the test code as a diff.
## Proof
Two results with commands: without the fix (failing, with the assertion message) and with the fix (passing). If the bug is not fixed yet, the failing run only.
## Notes
Anything that limits the test, such as a bug that is only observable end to end, or "None".
</output_format>
````

---

<a id="add-characterization-tests"></a>

## Add characterization tests to legacy code

`add-characterization-tests` · prompt · Testing · https://hermes-ide.com/prompts/add-characterization-tests

Pins down what untested legacy code does today with characterization and golden-master tests, bugs included, so it can be changed safely. Use before refactoring or modifying code with no tests.

````markdown
<context>
A characterization test records what the code actually does, not what it should do. It is a safety net for a later change: if a refactor alters any output, a test fails. That means the tests must pin current behaviour exactly, including odd and probably wrong behaviour, and must fail when the behaviour changes. Tests that only check "no exception" or that assert what the author guessed the code does give false confidence.
</context>

<task>
Write characterization tests for:
[CODE]


1. Find the entry points (from the list above, or from callers in the repository) and test through the highest-level one that is practical to call. Avoid testing private helpers that a refactor will move.
2. Find the seams that make the code nondeterministic or hard to call: current time, randomness, generated ids, environment, file system, network, database, global state. For each, choose the least invasive way to control it: an existing parameter or injection point first, then a test double at the module boundary, then a minimal seam (extract a parameter with the current value as its default). Name any production change you need; keep it behaviour-preserving.
3. Choose inputs that exercise every branch you can see: typical values, boundaries, empty and missing values, error paths, and combinations of flags. Read the conditionals to derive them.
4. Capture current outputs:
   - for small outputs, assert exact values;
   - for large or structured outputs (reports, HTML, JSON, files), write a golden-master or approval test that stores the output in a snapshot file, with scrubbers that normalise timestamps, ids and unordered collections so the snapshot is stable;
   - record side effects too: calls to collaborators, rows written, messages sent, exceptions raised.
   Derive expected values by running the code where you can. If you cannot run it, derive them by tracing the code and mark those tests "traced, confirm on first run".
5. Check the net catches change: for each important branch, describe a small mutation (flip a comparison, drop a line) and confirm a test would fail. Add inputs where none would.
</task>

<constraints>
- Do not fix bugs. Pin the current behaviour and list it under "Suspicious behaviour", with the test name, so a human decides later.
- Do not refactor production code beyond the minimal seams named in step 2.
- Name tests by behaviour (`returns_zero_discount_when_cart_empty`), not by number.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
</constraints>

<output_format>
## Behaviour inventory
Table: Entry point | Input class | Current output or side effect.
## Seams
Bullets: the nondeterminism or dependency, and how the tests control it (including any production change).
## Tests
The complete test file or files, with snapshot files if any.
## Suspicious behaviour
Table: Behaviour | Test that pins it | Why it looks wrong. Or "None".
## Coverage and gaps
Branches covered, branches not covered and why, and the mutations you checked.
</output_format>
````

---

<a id="fill-test-gaps"></a>

## Find and fill the riskiest test gaps

`fill-test-gaps` · prompt · Testing · https://hermes-ide.com/prompts/fill-test-gaps

Finds untested behaviour that matters most, ranked by risk rather than coverage percentage, and writes tests for the top gaps. Use when a module feels under-tested or before a risky change.

````markdown
<context>
Coverage percentage measures which lines ran, not which behaviours are checked. A module can show 90% coverage while its error handling, money arithmetic and permission checks are never asserted. The useful question is which untested behaviour would hurt most if it broke.
</context>

<task>
Find the riskiest test gaps in [SCOPE] and fill up to 5 of them.
1. Map the behaviours in scope: public functions, endpoints, state transitions, error paths, validations, permission checks.
2. Map the existing tests to those behaviours. A behaviour counts as covered only if a test asserts its result. Lines that merely run do not count.
3. Rank each uncovered behaviour by impact (money, data loss, security, user-visible failure) times likelihood (complex logic, recent churn in `git log`, past bugs, many callers).
4. Write tests for the top 5 gaps, following the project's existing test conventions. Each test must assert a specific result.
5. Run them. A test that fails on current code may have found a bug: keep it, mark it as expected to fail or skipped with a clear reason using the framework's mechanism, and report it. Do not change production code.
</task>

<constraints>
- Rank by risk, not by how easy a test is to write.
- Do not write tests whose only purpose is to raise coverage, such as tests that call code without asserting a result, or tests of trivial getters.
- Cite `path:line` for every gap.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Gaps
A table, highest risk first: # | Behaviour | Where | Why it is risky | Filled (yes or no).
## Tests
The new tests as a diff.
## Run
The command and its result. List any test that exposed a bug, with input, expected and actual.
## Remaining gaps
The gaps you did not fill, one line each, or "None".
</output_format>
````

---

<a id="fix-flaky-test"></a>

## Fix a flaky test

`fix-flaky-test` · prompt · Testing · https://hermes-ide.com/prompts/fix-flaky-test

Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.

````markdown
<context>
A flaky test passes and fails on the same code. Retries and longer timeouts hide the defect and teach the team to ignore red builds, so the goal is the cause, not a green run. Sometimes the flakiness is in the product code rather than the test, and then it is a real bug that users can hit.
</context>

<task>
Investigate [TEST].
1. Read the test, its fixtures and setup, and the code it exercises before running anything.
2. List the sources of nondeterminism you can see:
   - time: the current date or time, time zones, timers, timeouts that are too tight;
   - randomness: random data, unseeded generators, generated ids;
   - ordering: unordered collections, query results without ORDER BY, parallel tests, test order;
   - shared state: globals, singletons, caches, databases, files or ports used by other tests;
   - concurrency: unawaited promises, background work, sleeps used for synchronisation;
   - the outside world: network, external services, environment variables, locale.
3. Reproduce the failure: run the test repeatedly, in random order, in parallel, or alongside the tests that run before it in CI. Report how often it fails.
4. Fix the cause: wait on the condition instead of a duration, inject the clock or the seed, isolate the state, sort before comparing. If the race is in the product code, fix it there and say so.
5. Run the test enough times to show the failure is gone, using the same method that reproduced it.
</task>

<constraints>
- Never add retries, sleeps or longer timeouts as the fix.
- Never delete, skip or quarantine the test as the fix. If quarantine is needed while the fix lands, say so separately.
- If you cannot reproduce the failure, say so, report the most likely causes ranked with evidence, and do not claim a fix.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Cause
One paragraph: the nondeterminism and how it makes the test fail. Say whether it is in the test or in the product code.
## Fix
The diff, then one sentence on why it removes the cause.
## Evidence
Runs before and after, with the method used and failure counts (for example "7 of 200 failed before, 0 of 200 after").
</output_format>
````

---

<a id="review-test-quality"></a>

## Review test quality

`review-test-quality` · prompt · Testing · https://hermes-ide.com/prompts/review-test-quality

Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.

````markdown
<context>
A test earns its maintenance cost only if it fails when the behaviour it covers breaks and passes otherwise. Many tests do neither: they assert that a result is "not null", verify that a mock was called with whatever the mock returned, pass because an async assertion never ran, break when an internal method is renamed, depend on the order the suite runs in, or sleep and hope. Coverage numbers do not reveal any of this. The quickest way to judge a test is to ask which plausible bug in the code under test it would catch.
</context>

<task>
Review these tests:

<tests>
[TESTS]
</tests>

1. For each test, state in one line the behaviour it claims to check, judged from its name and body.
2. Look for tests that cannot fail: no assertion; assertions inside callbacks, loops or branches that may never run; un-awaited promises or async assertions; exceptions swallowed by `try`/`catch`; expected values computed with the same logic as the code; and comparisons of a mock's return value with itself.
3. Look for weak assertions: checking only existence, type, length or "truthy"; large snapshots nobody reads; asserting a subset when the whole result matters; and error tests that accept any exception instead of the specific one.
4. Look for over-mocking: mocking the unit under test or its pure collaborators, mocking types the project does not own instead of wrapping them, asserting call sequences instead of outcomes, and mocks whose behaviour differs from the real dependency (say how).
5. Look for hidden coupling: shared mutable fixtures, order dependence, global state, tests of private methods or internal structure, and one test covering several behaviours so a failure does not say what broke.
6. Look for nondeterminism: sleeps and fixed timeouts, real clocks and time zones, randomness without a seed, network or file-system dependence, unordered collections compared as ordered, concurrency without synchronisation, and locale-dependent formatting.
7. Mutation check: for the most important tests, name two or three small, realistic bugs in the code under test (an off-by-one, a flipped condition, a missing null check, a dropped field) and say whether each test would catch them. If the code under test was not provided, say what you infer and mark it as an inference.
8. Rewrite each problem test in the same framework and style, keeping its intent, so that it fails for the bug it should catch.

If the tests are fine, say so plainly and do not invent problems.
</task>

<constraints>
- Every finding cites the test name and line, the smell, the concrete bug it lets through or the false failure it causes, and the fix.
- Do not comment on naming or formatting unless it hides what is tested.
- Rewrites stay in the project's framework, helpers and conventions; no new test libraries unless one is clearly needed, and then say why.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Verdict
One line: solid | usable with fixes | gives false confidence. Then the main reason.

## Findings
Numbered, most harmful first. Each: `test name:line` - smell - what it lets through or breaks on - fix.

## Bugs these tests would miss
Table: plausible bug | caught? | by which test, or which test should catch it.

## Rewrites
Code blocks with the corrected tests, one per finding that needs code.
</output_format>
````

---

<a id="test-engineer"></a>

## Test engineer

`test-engineer` · persona · Testing · https://hermes-ide.com/prompts/test-engineer

Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.

````markdown
From now on, work as this persona: Test engineer.

You are a test engineer. You judge a test by one question: would it fail if the behaviour it describes broke? A suite that is green by default proves nothing, so you make sure each test can fail.

How you work:
- You start from behaviour: what the code promises its callers, including errors and limits. You read the code to find the branches, then test through the public interface, not the internals.
- You choose the cheapest level that can prove the behaviour: a unit test before an integration test before an end-to-end test. You go higher only when the risk lives in the wiring.
- You follow the project's existing test conventions, such as framework, layout, naming and fixtures, rather than introducing new ones.
- You watch every new test fail once, by breaking the behaviour or inverting the assertion, before you trust it.
- You treat flakiness as a defect with a cause: time, randomness, ordering, shared state, concurrency or the network.

What you flag:
- Tests that cannot fail: no assertion, assertions on mocks only, `expect(x).toBeTruthy()` where a value is known, snapshots nobody reads.
- Over-mocking: mocks of the code under test or of plain data, and tests that break on every refactor.
- Shared state between tests, order dependence, and real clocks, network or randomness inside unit tests.
- Retries, sleeps and skipped tests used to make a build green.
- Missing boundaries: empty, one, many, maximum, invalid, duplicate, Unicode, time zones, money rounding.

Your habits:
- You name tests after behaviour, so a failure message reads as a sentence about what broke.
- You keep one reason to fail per test and arrange, act and assert in that order.
- You report bugs you find instead of quietly changing production code to make a test pass.
- You report the command you ran and its real result.
````

---

<a id="test-writing-rules"></a>

## Test-writing rules

`test-writing-rules` · rule · Testing · https://hermes-ide.com/prompts/test-writing-rules

Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.

````markdown
Follow these rules for the rest of this conversation.

Apply these rules to files matching: `**/*.test.*`, `**/*.spec.*`, `**/*_test.*`, `**/test_*.py`.

When you write or change tests in this project:

**What to test**
- Test observable behaviour through the public interface: return values, state others can see, emitted events, HTTP responses, rendered output. Do not assert on private functions, internal call order or intermediate variables.
- Cover the cases that break code: empty input, a single item, boundaries, invalid input, error paths and concurrency where it applies, not just the happy path.
- Every bug fix comes with a test that fails without the fix.

**Shape**
- Each test checks one behaviour and has one reason to fail. Several assertions are fine when they describe the same behaviour.
- Name tests after the behaviour and the condition, such as "returns 404 when the order does not exist", not "test_get_2".
- Structure tests as arrange, act, assert, and set up only the data the test needs, using builders or factories with clear defaults.
- Make assertions specific: exact values, specific error types and messages. Avoid snapshot assertions of large output unless someone reviews the snapshot.

**Determinism**
- Never use sleeps to wait for something. Wait on the condition or event with a timeout, or use the framework's async utilities.
- Control time with a fake clock, randomness with a fixed seed, and time zone and locale explicitly. Never depend on the current date.
- Tests must not depend on execution order or on state left by other tests. Clean up files, records and global state, and give each test its own data.
- No real network calls to third parties in unit tests.

**Test doubles**
- Mock or fake only at the boundaries you do not own or cannot run cheaply: network, clock, file system, third-party services. Do not mock the unit under test or its internal collaborators.
- Prefer simple fakes and stubs to mocks with strict call expectations, which break on harmless refactors.

**Integrity**
- Follow the project's test framework, file layout and helpers. Do not add a new test library without asking.
- Keep unit tests fast, and mark slow or integration tests the way the project does.
- Run the tests you wrote and report the real result. If you could not run them, say so.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
````

---

<a id="write-e2e-test"></a>

## Write a resilient end-to-end test

`write-e2e-test` · prompt · Testing · https://hermes-ide.com/prompts/write-e2e-test

Writes an end-to-end browser test for a user flow with role-based locators, auto-waiting assertions and isolated test data, never fixed sleeps. Use when adding UI coverage for a critical path.

````markdown
<context>
End-to-end tests are the most expensive tests to keep green. They become flaky when they locate elements by CSS structure or generated class names, wait with fixed sleeps, share data between runs, or assert on things a user never sees. A resilient test finds elements the way a user or assistive technology does (role and accessible name, label, visible text), waits on conditions instead of time, owns its data, and checks the outcome the user cares about.
</context>

<task>
Write a playwright test for this flow:
[FLOW]

1. Restate the flow as numbered user actions, each with the observable outcome that proves it worked. If a step's expected outcome is not stated, ask for it or mark your assumption.
2. If you have the repository, read the relevant pages or components and any existing e2e setup (config, fixtures, page objects, auth helpers, test-data factories) and reuse them. Match the existing style. If the project already uses a different end-to-end framework than playwright, say so and ask which to use before writing.
3. Locators, in this order of preference:
   - Playwright: `getByRole` with name, then `getByLabel`, `getByPlaceholder`, `getByText`, then `getByTestId` as a last resort.
   - Cypress: Testing Library queries (`findByRole`, `findByLabelText`) if the project has them, otherwise `cy.contains` scoped to a container, then `data-testid`/`data-cy`.
   - Selenium: accessible attributes, labels and visible text via stable XPath or CSS on `data-testid`; never absolute XPath.
   Never use generated class names, nth-child chains or DOM position.
4. Waiting: use auto-retrying, web-first assertions (Playwright `expect(locator).toBeVisible()`/`toHaveText()`, Cypress `should`, Selenium `WebDriverWait` with expected conditions). Wait for a specific network response or UI state when an action triggers one. No `waitForTimeout`, `cy.wait(<ms>)` or `Thread.sleep`.
5. Isolation: create the data the test needs through an API, fixture or seed helper, with unique values per run, and clean it up or make it disposable. Log in through a stored session or API helper rather than the login form, unless login is the flow under test.
6. Assert the user-visible outcome at each checkpoint, plus one durable side effect if it matters (the saved record, the confirmation email stub), not implementation details.
</task>

<constraints>
- Do not invent selectors, routes or accessible names you have not seen. When the page source is not available, write the most likely role and name, and list each one under "Assumptions to verify".
- One flow per test. Keep the test independent of test order.
- If a step depends on a third-party service (payments, email, maps), stub it at the network layer and say so.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Test plan
Numbered steps: action, then expected outcome.
## Test
The complete test file in one code block, including setup and teardown helpers it needs.
## Assumptions to verify
Bullets: each selector, route or data assumption you could not confirm. Or "None".
## How to run
The command to run this one test headed and headless, and how to see the trace or screenshots on failure.
</output_format>
````

---

<a id="write-test-plan"></a>

## Write a test plan

`write-test-plan` · prompt · Testing · https://hermes-ide.com/prompts/write-test-plan

Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.

````markdown
<context>
A test plan is useful when it tells a team where to spend limited testing time and when to stop. Plans that list every possible test case get skimmed and ignored; plans with no risk analysis spread effort evenly, so the payment edge case gets the same attention as a label change. A good plan ranks risks, picks the cheapest test level that addresses each one, names what automation will not catch (usability, unusual data, real devices, integrations with real third parties, migration of existing data), and defines exit criteria that someone can actually check on release day.
</context>

<task>
Write a test plan for:
[FEATURE]

1. Define the scope: what is being tested (functions, platforms, user types, integrations) and what is explicitly out of scope, with the reason.
2. Identify risks: combine the known worries with what the feature implies (money, permissions, data migration, concurrency, third parties, performance, accessibility, localisation, backward compatibility, feature-flag states). Rate each by likelihood and impact, and rank them.
3. For each top risk, choose the test level that addresses it most cheaply (unit, integration, contract, end-to-end, manual exploratory, non-functional), say whether existing automation already covers it, and what new tests are needed. Name the gaps automation will not close.
4. Write exploratory charters for the manual work, in the form "Explore <area> with <resources or data> to discover <kind of problem>", each time-boxed, covering what scripted tests miss: unexpected sequences, interrupted flows, odd data, permissions, different devices and assistive technology.
5. Specify environments and test data: which environment, which configuration and feature-flag states, accounts and roles needed, data volume and edge records, third-party sandboxes, and how data is created and reset. No real personal data.
6. Define entry criteria (what must be true before testing starts) and exit criteria that are checkable: no open critical or high defects, the named risks covered, automated suites green, performance within stated limits, and an explicit decision on known issues. Include a rollback or flag-off check if the release can be reverted.
7. Lay out the schedule against the release date, with owners as roles, and say what to cut first if time runs short (lowest-ranked risks), so the trade-off is visible rather than accidental.

If the feature description is too thin to identify risks (no behaviour, users or integrations), ask for the spec or acceptance criteria and stop.
</task>

<constraints>
- Rank everything by risk. Do not list low-value test cases to look thorough.
- Do not duplicate what existing automation already covers; reference it instead.
- Exit criteria must be measurable or a named decision, never "sufficient testing done".
- Do not invent dates, people or metrics; use roles and placeholders.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Scope
In scope and out of scope, as two short lists.

## Risks
Table: # | risk | likelihood | impact | priority.

## Test approach
Table: risk # | test level | covered by existing automation? | new tests needed.

## Environments and data
Bullets.

## Manual and exploratory testing
Numbered charters with time boxes, plus any must-do manual checks.

## Entry and exit criteria
Two checklists.

## Schedule and owners
Table: activity | owner (role) | when. Then "If time runs short, cut:" in priority order.

## Open questions
Only the ones that change the plan.
</output_format>
````

---

<a id="write-contract-tests"></a>

## Write consumer-driven contract tests

`write-contract-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-contract-tests

Writes consumer-driven contract tests between two services and the CI gate that runs them, so a breaking API change fails before deploy. Use when services that call each other ship independently.

````markdown
<context>
Contract tests catch the integration bug where each service passes its own tests but the provider renames a field, tightens validation or changes a status code that a consumer depends on. In consumer-driven contracts the consumer records only what it actually sends and reads, so the provider is free to change everything else. Contracts that copy whole responses with exact values are brittle and block harmless changes; contracts that are never verified against the real provider, or never gate a deploy, catch nothing.
</context>

<task>
Write contract tests using the pact approach.

Consumer:
[CONSUMER]

Provider API:
[PROVIDER_API]

1. List each interaction the consumer really uses: method, path, query, headers that matter, request body fields, the response status and only the response fields the consumer reads. Trace field reads in the consumer code; do not include fields it ignores. Include the error responses the consumer handles (404, 409, 422 and so on).
2. For each interaction, name the provider state it needs ("order 42 exists and is paid").
3. Consumer side: write tests that exercise the real client code against the contract mock and check the client's own parsing, not just the mock. Use type and format matchers (like-type, regex, each-like with a minimum) instead of literal values, except where the exact value is the contract (enums, status codes).
4. Provider side: write the verification that replays the contract against the running provider, with a state handler per provider state that sets up data through the provider's own code or test fixtures.
5. CI: show how the contract is published (with consumer version and branch), how the provider verifies on every build, and the pre-deploy check that blocks a deploy when the deployed counterpart's contract is not verified (for Pact, a broker or PactFlow with `can-i-deploy --to-environment`, `record-deployment` after each deploy so the broker knows what runs where, and a contract-changed webhook that triggers provider verification). For openapi-schema, validate consumer mocks against the spec and provider responses against the spec, and fail on spec drift. For custom, store fixtures in one place both builds read, and version them.
6. Show one concrete breaking change (a renamed field, say) and which check fails.
</task>

<constraints>
- Do not invent endpoints, fields or status codes that are not in the inputs. If the provider API and the consumer disagree, report the mismatch as a finding instead of choosing one.
- Contracts are not functional tests: do not assert business rules of the provider beyond the shape and semantics the consumer relies on.
- Use the client library and test runner the consumer already uses. Name any package to install with its ecosystem.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Interactions
Table: Interaction | Request | Response fields used | Provider state.
## Consumer tests
Complete test file(s) in code blocks.
## Provider verification
Complete verification test with state handlers.
## CI gate
The pipeline steps (as config or a numbered list) for publish, verify and the pre-deploy check, and the breaking-change example.
## What this does not catch
Bullets: behaviour contract tests miss here (performance, auth flows, data semantics) and what covers it instead. Mismatches found between consumer and provider go first, if any.
</output_format>
````

---

<a id="write-integration-tests"></a>

## Write integration tests with real dependencies

`write-integration-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-integration-tests

Writes integration tests that run against real dependencies such as databases and queues in containers, with fixtures, isolation between tests and cleanup. Use when mocks hide bugs at the boundary.

````markdown
<context>
Integration tests exist to catch what mocks cannot: SQL that only fails on the real engine, transaction and locking behaviour, migrations, serialisation across a queue, unique constraints, time zones and encodings. They become a burden when they share state and fail in random order, sleep instead of waiting, start a fresh container per test and take twenty minutes, or test the dependency rather than the code. Good integration tests start each dependency once per run, give every test its own data, wait on conditions, and assert on observable outcomes.
</context>

<task>
Write integration tests for:
<code>
[CODE]
</code>

1. Read the code and the project's existing test setup (framework, runner, folders, helpers, migrations, CI config). Follow what exists. If you cannot see the code or the dependency versions, ask once for what is missing and stop.
2. Write a short test plan: the behaviours that cross a real boundary (queries with filtering and ordering, constraint violations, transactions and rollbacks, concurrent updates, message publish and consume, retries and dead-lettering, cache expiry), each with the outcome to assert. Leave pure logic to unit tests.
3. Set up dependencies in containers, preferring the Testcontainers library for the language, or a compose file the test run starts. Pin image versions to match production. Start each container once per test run or suite, not per test. Apply the real schema migrations, not a hand-written schema.
4. Isolate tests. Pick the cheapest strategy that is correct and say why: a transaction per test rolled back at the end (not valid when the code under test commits or uses several connections), a unique schema, database, queue or key prefix per test or worker, or truncating tables between tests. Make tests safe to run in parallel or mark them serial.
5. Build data with small factories or builders that set only the fields a test cares about. No shared mutable fixtures.
6. Wait on conditions with a timeout (poll until the message is consumed, up to a few seconds); never fixed sleeps.
7. Clean up containers, connections and temporary resources even when a test fails.
8. Run the tests and report the real result. If you cannot run them (no container runtime), say so plainly.
</task>

<constraints>
- Use the real dependency for the behaviour under test; mock only external third parties you do not control, and say which.
- Never point tests at a shared or production environment, and never read real credentials. Use container-generated connection settings.
- Assert on outcomes (rows, messages, responses), not on which internal functions were called.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Test plan
A table: behaviour, dependency, assertion.
## Setup
The container or compose setup and shared fixtures, as code blocks with file paths.
## Tests
The test files, as code blocks with file paths.
## How to run
Commands for local runs and the CI job change, plus the result of running them.
## Notes
Isolation strategy chosen and why, expected runtime, and anything that could make the tests flaky.
</output_format>
````

---

<a id="write-property-based-tests"></a>

## Write property-based tests

`write-property-based-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-property-based-tests

Finds the invariants a function must keep and writes property-based tests with generators that shrink well. Use when example-based tests miss edge cases in parsers, encoders or pure logic.

````markdown
<context>
Property-based tests state a rule that must hold for every valid input and let a generator search for a counterexample, then shrink it to the smallest failing case. They find the bugs example tests miss, but only when the property is genuinely true of the specification (not a restatement of the implementation) and the generators produce valid, varied, shrinkable inputs. A property that re-implements the function proves nothing; a generator that filters away 90% of its draws is slow and shrinks badly.
</context>

<task>
Write property-based tests for:
[CODE]

Library:  If no library is named, detect it from the project's manifests and existing tests (Hypothesis for Python, fast-check for JavaScript and TypeScript, proptest for Rust, jqwik for Java, FsCheck for .NET, rapid for Go; the standard library's testing/quick is frozen and shrinks nothing). If none is installed, pick the standard one for the language and say how to add it.

1. Read the code and state its contract: valid inputs, outputs, errors it may raise, and side effects. If the contract is ambiguous (for example, what happens on empty input), ask or state the assumption you test against.
2. Find candidate properties, preferring these patterns:
   - round-trip: decode(encode(x)) == x, parse(print(x)) == x;
   - invariants: output is sorted, length preserved, total conserved, no duplicates, within bounds;
   - idempotence: f(f(x)) == f(x);
   - oracle or model: agrees with a simpler, obviously correct implementation or an in-memory model of a stateful system;
   - metamorphic: a known change to the input causes a predictable change to the output;
   - algebraic: commutativity, associativity, identity elements where the domain promises them;
   - robustness: never crashes or hangs on any input of the right type, and fails only with documented errors.
   Keep only properties that follow from the contract. Discard any that just mirror the implementation.
3. Build generators from the domain, not from raw types: construct valid values directly (map, compose, build strategies) instead of generating anything and filtering. Include the edge values the type allows: empty, single element, zero, negative, maximum sizes, Unicode beyond ASCII, NaN and infinities for floats where relevant. Bound sizes so a run stays fast.
4. Write the tests in the project's style and test runner. Make failures reproducible: rely on the library's seed reporting and example database or replay, and add any shrunk counterexample you discover as an explicit regression example.
5. If you can run the tests, do so and report the result. If a property fails, report the minimal counterexample and whether the bug is in the code or in your property. Do not change the code under test.
</task>

<constraints>
- Every property must name the contract clause it checks. No property may call the function under test to compute its own expected value.
- Avoid filter or assume calls that reject more than a small fraction of draws; restructure the generator instead.
- Keep default example counts unless there is a reason to change them, and say why if you do.
- Do not fix bugs you find; report them.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
</constraints>

<output_format>
## Properties
Table: Property | Pattern | Contract clause it checks.
## Generators
One line per generator: what it builds and which edge values it covers.
## Tests
The complete test file in one code block, with imports.
## Counterexamples
Shrunk failing inputs with a one-line diagnosis each, or "None found" with the number of examples run. If you could not run the tests, say so.
## How to run
The exact command, including how to replay a failure from its seed.
</output_format>
````

---

<a id="write-unit-tests"></a>

## Write unit tests

`write-unit-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-unit-tests

Writes unit tests that pin a unit's behaviour, covering boundaries, errors and edge inputs in the project's own test style, and proves each test can fail. Use for new or untested code.

````markdown
<context>
Good unit tests describe what a unit does, not how it does it. They fail when behaviour breaks and keep passing through refactors. Tests that mirror the implementation, mock everything, or assert only that no exception was thrown add maintenance cost without catching bugs.
</context>

<task>
Write unit tests for [TARGET].
1. Read the target and its callers to learn its contract: inputs, outputs, side effects, errors. Read two or three existing test files to learn the project's conventions (framework, file location, naming, fixtures, assertion style) and follow them.
2. List the behaviours to cover before writing any test:
   - the main cases;
   - boundaries: empty, one element, maximum, zero, negative, off-by-one limits;
   - invalid input and every error path the code defines;
   - inputs that often break code: null or missing values, duplicates, Unicode, very large values, time zones and dates, floating-point amounts.
3. Write one test per behaviour, through the unit's public interface. Name each test after the behaviour (`returns empty list when no orders match`), not after the method.
4. Use fakes or mocks only at real boundaries: network, clock, file system, randomness, other services. Do not mock the code under test or plain data objects.
5. Run the tests. For each new test, confirm it can fail: break the behaviour temporarily or invert the assertion, watch it fail, then restore it.
</task>

<constraints>
- Do not change production code. If the code is hard to test, or you find a bug, report it under "Not covered" with the failing input and leave the code alone.
- Each test asserts specific values, not only that something is truthy or that no error was thrown.
- Keep tests independent: no shared mutable state and no dependence on run order.
- No snapshot tests unless the project already uses them for this kind of output.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Behaviours
A table: Behaviour | Test name | Kind (main, boundary, error, edge).
## Tests
The new or changed test files as a diff.
## Run
The command you ran and its result, plus how you confirmed the tests can fail.
## Not covered
Behaviours you did not test and why, and any bugs found (input, expected, actual). Or "None".
</output_format>
````
