hermes

Debug a race condition

Diagnoses intermittent concurrency bugs by mapping shared state and the interleavings that break it, adds targeted instrumentation and proposes a fix. Use for bugs seen only under load.

context

Race conditions are bugs in ordering: two or more units of execution touch the same state, and some interleaving of their steps breaks an invariant. They hide from debuggers and print statements because observing them changes the timing. The reliable way in is to reason from the shared state and the possible interleavings, form specific hypotheses, then make the bad interleaving more likely on purpose and prove it with evidence. Sleeps, retries and "add a lock somewhere" usually move the bug rather than remove it.

task

Diagnose this concurrency bug.

Code:

Symptoms:

Only if [RUNTIME] is given: Runtime and concurrency model: If the runtime or the concurrency model is not clear from the code, ask before going further, because the answer changes which interleavings are possible.

  1. Map the concurrency: list each unit that runs concurrently (threads, goroutines, async tasks, workers, processes, app instances, cron jobs) and each piece of shared state (in-memory fields, caches, globals, files, database rows, queues, external resources). For each piece, list every read and write with its location and the synchronisation that protects it, if any.
  2. Name the invariant that the symptom shows is broken (for example, "an order is charged at most once").
  3. Enumerate candidate interleavings that break it. Check at least: check-then-act and read-modify-write without atomicity; lost updates in the database under the actual isolation level; publication without a happens-before edge (unsafe lazy init, non-volatile flags); iterating a collection while it is modified; await points that split a critical section in single-threaded async code; lock ordering that can deadlock; time-of-check to time-of-use on files or external state; duplicate delivery from retries or at-least-once queues. Write each candidate as a step-by-step timeline of A and B.
  4. Rank candidates by how well they explain every symptom (frequency, load dependence, the exact wrong value). Drop those that contradict the evidence.
  5. Propose instrumentation that can confirm or rule out the top candidates without hiding the bug: log lines with unit id, monotonic timestamp and a sequence or version number at each access; the runtime's race detector or concurrency checker if one exists for this runtime; a stress test that runs the operation concurrently many times, with injected delays or yields at the suspected gap to widen the window.
  6. Propose the fix that removes the bad interleaving at its root, preferring in order: removing the sharing, making the operation atomic (a single atomic op, a conditional update, a unique constraint, a transaction at the right isolation level, optimistic locking with a version), then a lock with a documented scope and order. Make operations idempotent where duplicates are possible.
  7. Define how to verify: the stress test fails before the fix at a measured rate and passes after many runs.
constraints
  • Never propose sleeps, retries or longer timeouts as the fix.
  • Do not claim a root cause is confirmed until the evidence from step 5 confirms it; until then, call it the leading hypothesis.
  • Keep the fix as small as the root cause allows, and state what it costs (contention, throughput, latency).
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Shared state

Table: State | Readers and writers (location) | Protection.

Candidate interleavings

Ranked. For each: the broken invariant, a two-column timeline (A | B), and how well it explains the symptoms.

Instrumentation

What to add or run, and the result that would confirm or rule out each top candidate.

Fix

The diff for the leading candidate, and why it removes the interleaving.

Verification

The stress or race-detector test, how many runs, and the before and after failure rates to expect.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Debugging
level
Expert
made for
Software engineer, Backend engineer
needs
repo-read
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install debug-race-condition --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill debug-race-condition -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Debugging
PersonaDebugging

Debugger

Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.

debugger
PersonaImplementation

Backend engineer

Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.

backend-engineer
PromptDebugging

Debug a failing network request

Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.

debug-network-request
PromptDebugging

Debug a production-only bug

Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.

debug-production-only-bug
PromptDebugging

Bisect a regression

Finds the commit or input that introduced a regression by writing an automated good/bad check first, then bisecting. Use when something that used to work is broken and the cause is unclear.

bisect-regression
WorkflowDebugging

Bugfix track

Takes a bug from report to reproduction, root cause, regression test, minimal fix and a verified pull request, stopping for approval between steps. Use for any bug worth fixing properly.

bugfix-track