hermes

Find and fill the riskiest test gaps

Finds untested behaviour that matters most, ranked by risk rather than coverage percentage, and writes tests for the top gaps. Use when a module feels under-tested or before a risky change.

context

Coverage percentage measures which lines ran, not which behaviours are checked. A module can show 90% coverage while its error handling, money arithmetic and permission checks are never asserted. The useful question is which untested behaviour would hurt most if it broke.

task

Find the riskiest test gaps in and fill up to of them. Only if [COVERAGE_REPORT] is given: Coverage data:

  1. Map the behaviours in scope: public functions, endpoints, state transitions, error paths, validations, permission checks.
  2. Map the existing tests to those behaviours. A behaviour counts as covered only if a test asserts its result. Lines that merely run do not count.
  3. Rank each uncovered behaviour by impact (money, data loss, security, user-visible failure) times likelihood (complex logic, recent churn in git log, past bugs, many callers).
  4. Write tests for the top gaps, following the project's existing test conventions. Each test must assert a specific result.
  5. Run them. A test that fails on current code may have found a bug: keep it, mark it as expected to fail or skipped with a clear reason using the framework's mechanism, and report it. Do not change production code.
constraints
  • Rank by risk, not by how easy a test is to write.
  • Do not write tests whose only purpose is to raise coverage, such as tests that call code without asserting a result, or tests of trivial getters.
  • Cite path:line for every gap.
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
  • Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
  • If a test looks wrong, explain why and ask before changing it.
  • Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
  • If you could not run a check, say so plainly and say which one.
output format

Gaps

A table, highest risk first: # | Behaviour | Where | Why it is risky | Filled (yes or no).

Tests

The new tests as a diff.

Run

The command and its result. List any test that exposed a bug, with input, expected and actual.

Remaining gaps

The gaps you did not fill, one line each, or "None".

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Testing
level
Intermediate
made for
Software engineer, QA / test engineer, Tech lead / staff engineer
needs
repo-read, file-write, shell
risk
runs-commands
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install fill-test-gaps --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill fill-test-gaps -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Testing
PersonaTesting

Test engineer

Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.

test-engineer
PromptTesting

Write unit tests

Writes unit tests that pin a unit's behaviour, covering boundaries, errors and edge inputs in the project's own test style, and proves each test can fail. Use for new or untested code.

write-unit-tests
PromptTesting

Review test quality

Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.

review-test-quality
RuleTesting

Test-writing rules

Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.

test-writing-rules
PromptTesting

Write a test plan

Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.

write-test-plan
PromptTesting

Add characterization tests to legacy code

Pins down what untested legacy code does today with characterization and golden-master tests, bugs included, so it can be changed safely. Use before refactoring or modifying code with no tests.

add-characterization-tests