hermes

Fix a flaky test

Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.

context

A flaky test passes and fails on the same code. Retries and longer timeouts hide the defect and teach the team to ignore red builds, so the goal is the cause, not a green run. Sometimes the flakiness is in the product code rather than the test, and then it is a real bug that users can hit.

task

Investigate . Only if [FAILURE_LOG] is given: Start from this failing output:

  1. Read the test, its fixtures and setup, and the code it exercises before running anything.
  2. List the sources of nondeterminism you can see:
  • time: the current date or time, time zones, timers, timeouts that are too tight;
  • randomness: random data, unseeded generators, generated ids;
  • ordering: unordered collections, query results without ORDER BY, parallel tests, test order;
  • shared state: globals, singletons, caches, databases, files or ports used by other tests;
  • concurrency: unawaited promises, background work, sleeps used for synchronisation;
  • the outside world: network, external services, environment variables, locale.
  1. Reproduce the failure: run the test repeatedly, in random order, in parallel, or alongside the tests that run before it in CI. Report how often it fails.
  2. Fix the cause: wait on the condition instead of a duration, inject the clock or the seed, isolate the state, sort before comparing. If the race is in the product code, fix it there and say so.
  3. Run the test enough times to show the failure is gone, using the same method that reproduced it.
constraints
  • Never add retries, sleeps or longer timeouts as the fix.
  • Never delete, skip or quarantine the test as the fix. If quarantine is needed while the fix lands, say so separately.
  • If you cannot reproduce the failure, say so, report the most likely causes ranked with evidence, and do not claim a fix.
  • Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
  • If a test looks wrong, explain why and ask before changing it.
  • Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
  • If you could not run a check, say so plainly and say which one.
output format

Cause

One paragraph: the nondeterminism and how it makes the test fail. Say whether it is in the test or in the product code.

Fix

The diff, then one sentence on why it removes the cause.

Evidence

Runs before and after, with the method used and failure counts (for example "7 of 200 failed before, 0 of 200 after").

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Testing
level
Intermediate
made for
Software engineer, QA / test engineer
needs
repo-read, file-write, shell
risk
runs-commands
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install fix-flaky-test --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill fix-flaky-test -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Testing
PersonaTesting

Test engineer

Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.

test-engineer
PromptDebugging

Triage a failing CI build

Finds the first real error in a failing CI log, classifies the failure as caused by the change, flaky, environment drift or already broken, and names the next action. Use when a pipeline turns red.

triage-failing-ci
PromptTesting

Review test quality

Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.

review-test-quality
RuleTesting

Test-writing rules

Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.

test-writing-rules
PromptTesting

Write a test plan

Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.

write-test-plan
PromptTesting

Add characterization tests to legacy code

Pins down what untested legacy code does today with characterization and golden-master tests, bugs included, so it can be changed safely. Use before refactoring or modifying code with no tests.

add-characterization-tests