Review test quality
Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.
A test earns its maintenance cost only if it fails when the behaviour it covers breaks and passes otherwise. Many tests do neither: they assert that a result is "not null", verify that a mock was called with whatever the mock returned, pass because an async assertion never ran, break when an internal method is renamed, depend on the order the suite runs in, or sleep and hope. Coverage numbers do not reveal any of this. The quickest way to judge a test is to ask which plausible bug in the code under test it would catch.
Review these testsOnly if [FRAMEWORK] is given: ():
Only if [CODE_UNDER_TEST] is given:
- For each test, state in one line the behaviour it claims to check, judged from its name and body.
- Look for tests that cannot fail: no assertion; assertions inside callbacks, loops or branches that may never run; un-awaited promises or async assertions; exceptions swallowed by
try/catch; expected values computed with the same logic as the code; and comparisons of a mock's return value with itself. - Look for weak assertions: checking only existence, type, length or "truthy"; large snapshots nobody reads; asserting a subset when the whole result matters; and error tests that accept any exception instead of the specific one.
- Look for over-mocking: mocking the unit under test or its pure collaborators, mocking types the project does not own instead of wrapping them, asserting call sequences instead of outcomes, and mocks whose behaviour differs from the real dependency (say how).
- Look for hidden coupling: shared mutable fixtures, order dependence, global state, tests of private methods or internal structure, and one test covering several behaviours so a failure does not say what broke.
- Look for nondeterminism: sleeps and fixed timeouts, real clocks and time zones, randomness without a seed, network or file-system dependence, unordered collections compared as ordered, concurrency without synchronisation, and locale-dependent formatting.
- Mutation check: for the most important tests, name two or three small, realistic bugs in the code under test (an off-by-one, a flipped condition, a missing null check, a dropped field) and say whether each test would catch them. If the code under test was not provided, say what you infer and mark it as an inference.
- Rewrite each problem test in the same framework and style, keeping its intent, so that it fails for the bug it should catch.
If the tests are fine, say so plainly and do not invent problems.
- Every finding cites the test name and line, the smell, the concrete bug it lets through or the false failure it causes, and the fix.
- Do not comment on naming or formatting unless it hides what is tested.
- Rewrites stay in the project's framework, helpers and conventions; no new test libraries unless one is clearly needed, and then say why.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
Verdict
One line: solid | usable with fixes | gives false confidence. Then the main reason.
Findings
Numbered, most harmful first. Each: test name:line - smell - what it lets through or breaks on - fix.
Bugs these tests would miss
Table: plausible bug | caught? | by which test, or which test should catch it.
Rewrites
Code blocks with the corrected tests, one per finding that needs code.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Testing
- level
- Intermediate
- made for
- Software engineer, QA / test engineer, Tech lead / staff engineer, Open-source maintainer
- needs
- repo-read
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install review-test-quality --target claude-codenpx skills add hermes-hq/hodios-dist --skill review-test-quality -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of TestingFix a flaky test
Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.
fix-flaky-testWrite unit tests
Writes unit tests that pin a unit's behaviour, covering boundaries, errors and edge inputs in the project's own test style, and proves each test can fail. Use for new or untested code.
write-unit-testsFind and fill the riskiest test gaps
Finds untested behaviour that matters most, ranked by risk rather than coverage percentage, and writes tests for the top gaps. Use when a module feels under-tested or before a risky change.
fill-test-gapsTest engineer
Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.
test-engineerTest-writing rules
Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.
test-writing-rulesWrite a test plan
Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.
write-test-plan