hermes

Review test quality

Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.

context

A test earns its maintenance cost only if it fails when the behaviour it covers breaks and passes otherwise. Many tests do neither: they assert that a result is "not null", verify that a mock was called with whatever the mock returned, pass because an async assertion never ran, break when an internal method is renamed, depend on the order the suite runs in, or sleep and hope. Coverage numbers do not reveal any of this. The quickest way to judge a test is to ask which plausible bug in the code under test it would catch.

task

Review these testsOnly if [FRAMEWORK] is given: ():

tests

Only if [CODE_UNDER_TEST] is given:

code under test

  1. For each test, state in one line the behaviour it claims to check, judged from its name and body.
  2. Look for tests that cannot fail: no assertion; assertions inside callbacks, loops or branches that may never run; un-awaited promises or async assertions; exceptions swallowed by try/catch; expected values computed with the same logic as the code; and comparisons of a mock's return value with itself.
  3. Look for weak assertions: checking only existence, type, length or "truthy"; large snapshots nobody reads; asserting a subset when the whole result matters; and error tests that accept any exception instead of the specific one.
  4. Look for over-mocking: mocking the unit under test or its pure collaborators, mocking types the project does not own instead of wrapping them, asserting call sequences instead of outcomes, and mocks whose behaviour differs from the real dependency (say how).
  5. Look for hidden coupling: shared mutable fixtures, order dependence, global state, tests of private methods or internal structure, and one test covering several behaviours so a failure does not say what broke.
  6. Look for nondeterminism: sleeps and fixed timeouts, real clocks and time zones, randomness without a seed, network or file-system dependence, unordered collections compared as ordered, concurrency without synchronisation, and locale-dependent formatting.
  7. Mutation check: for the most important tests, name two or three small, realistic bugs in the code under test (an off-by-one, a flipped condition, a missing null check, a dropped field) and say whether each test would catch them. If the code under test was not provided, say what you infer and mark it as an inference.
  8. Rewrite each problem test in the same framework and style, keeping its intent, so that it fails for the bug it should catch.

If the tests are fine, say so plainly and do not invent problems.

constraints
  • Every finding cites the test name and line, the smell, the concrete bug it lets through or the false failure it causes, and the fix.
  • Do not comment on naming or formatting unless it hides what is tested.
  • Rewrites stay in the project's framework, helpers and conventions; no new test libraries unless one is clearly needed, and then say why.
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
output format

Verdict

One line: solid | usable with fixes | gives false confidence. Then the main reason.

Findings

Numbered, most harmful first. Each: test name:line - smell - what it lets through or breaks on - fix.

Bugs these tests would miss

Table: plausible bug | caught? | by which test, or which test should catch it.

Rewrites

Code blocks with the corrected tests, one per finding that needs code.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Testing
level
Intermediate
made for
Software engineer, QA / test engineer, Tech lead / staff engineer, Open-source maintainer
needs
repo-read
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install review-test-quality --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill review-test-quality -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Testing
PromptTesting

Fix a flaky test

Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.

fix-flaky-test
PromptTesting

Write unit tests

Writes unit tests that pin a unit's behaviour, covering boundaries, errors and edge inputs in the project's own test style, and proves each test can fail. Use for new or untested code.

write-unit-tests
PromptTesting

Find and fill the riskiest test gaps

Finds untested behaviour that matters most, ranked by risk rather than coverage percentage, and writes tests for the top gaps. Use when a module feels under-tested or before a risky change.

fill-test-gaps
PersonaTesting

Test engineer

Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.

test-engineer
RuleTesting

Test-writing rules

Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.

test-writing-rules
PromptTesting

Write a test plan

Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.

write-test-plan