# Hodios paste pack: Debugging

Everything in Debugging from Hodios, the open prompt library by Hermes IDE: 11 entries, catalog 2026.1003.0.

Every entry is dedicated to the public domain under CC0 1.0. Copy, change and share them freely, no attribution needed.

Browse and search the library at https://hermes-ide.com/prompts

## How to use

Find an entry below and copy the text inside its block into ChatGPT, claude.ai or any chat. Replace each [PLACEHOLDER] with your own material. Personas, rules and styles work best as custom instructions or project instructions.

## Contents

- Debugging
  - [Bisect a regression](#bisect-regression) (prompt)
  - [Bugfix track](#bugfix-track) (workflow)
  - [Debug a failing network request](#debug-network-request) (prompt)
  - [Debug a mobile app crash](#debug-mobile-crash) (prompt)
  - [Debug a production-only bug](#debug-production-only-bug) (prompt)
  - [Debug a race condition](#debug-race-condition) (prompt)
  - [Debugger](#debugger) (persona)
  - [Explain a stack trace](#explain-stack-trace) (prompt)
  - [Find the root cause of a bug](#find-root-cause) (prompt)
  - [Triage a failing CI build](#triage-failing-ci) (prompt)
  - [Turn a bug report into a minimal reproduction](#reproduce-bug-report) (prompt)

---

<a id="bisect-regression"></a>

## Bisect a regression

`bisect-regression` · prompt · Debugging · https://hermes-ide.com/prompts/bisect-regression

Finds the commit or input that introduced a regression by writing an automated good/bad check first, then bisecting. Use when something that used to work is broken and the cause is unclear.

````markdown
<context>
Bisection finds the first bad commit in log2(n) steps, but only if every step is judged correctly. Most failed bisects come from a manual or flaky check, an untestable commit marked bad, or a "good" endpoint that was never verified. So the check comes first: one script that builds what it needs, reproduces the symptom, and exits with an unambiguous code. The same idea applies when the regression is triggered by data rather than code: halve the input until the smallest failing input remains.
</context>

<task>
Find what introduced this regression:
[REGRESSION]

Bad: HEAD. 

1. **Write the check.** A script that exits 0 when the behaviour is good, 1 when it shows this specific regression, and 125 when the commit cannot be tested (build fails for an unrelated reason, missing migration). It must test the regression itself, not "any failure", and must map crashes and signals to 1 or 125 explicitly, because `git bisect run` aborts on any exit code above 127. Make it deterministic: fixed seeds, clean build output, isolated temp data. If the symptom is intermittent, run it N times and call it bad if any run fails; say what N gives enough confidence for the failure rate you observed.
2. **Confirm the endpoints.** Run the check on the bad ref and the good ref and show the results. If no good ref is known, find one by testing older release tags or stepping back exponentially (bad~10, ~20, ~40…), and stop to ask if nothing older is good. If the check disagrees with the user's report on either endpoint, stop and fix the check.
3. **Decide what to bisect.** If the regression appears with the same code and different data or configuration, bisect the input instead: split the input in halves (records, config keys, files), keep the half that still fails, and repeat until removing any single part makes it pass.
4. **Run the bisect:** `git bisect start <bad> <good>`, then `git bisect run <check>`. Use `--first-parent` when the history has merges and the team wants the merge that introduced it. Note any skipped commits.
5. **Confirm the culprit.** Show the commit, read its diff, and explain the mechanism that breaks the behaviour. Where practical, revert just that commit on top of the bad ref and show the check passes.
6. End with `git bisect reset` and say which branch is checked out.
</task>

<constraints>
- Never mark a commit bad because it fails to build or fails for a different reason; that is a skip (exit 125).
- Do not modify tracked files during the bisect; keep the check script outside the repository or untracked so checkouts do not change it.
- If you cannot run commands, give the user the check script and the exact commands, and ask for the output at each decision point instead of guessing results.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Check
The script in a code block, and what each exit code means for this regression.
## Good and bad endpoints
The refs and the check result on each.
## Bisect
The exact commands, and the bisect log if you ran it.
## Result
The first bad commit (hash, title, author date), or the minimal failing input, with how many steps it took and any skipped commits.
## Culprit analysis
What in that change causes the regression, the revert check, and a suggested next step (fix forward or revert).
</output_format>
````

---

<a id="bugfix-track"></a>

## Bugfix track

`bugfix-track` · workflow · Debugging · https://hermes-ide.com/prompts/bugfix-track

Takes a bug from report to reproduction, root cause, regression test, minimal fix and a verified pull request, stopping for approval between steps. Use for any bug worth fixing properly.

````markdown
Fixes this bug properly, one approved step at a time:

<bug_report>
[BUG_REPORT]
</bug_report>

Severity: medium.

The order is fixed: reproduce it, find the root cause, write a test that fails because of the bug, make the smallest fix that turns the test green, then verify everything and prepare the pull request. Each step ends with a short report and stops for the developer's approval; later steps build on the approved findings instead of re-asking. Nothing is called fixed until a test that failed before the change passes after it and the rest of the suite still passes. If the severity is high or critical, the first step also says whether users need a mitigation now (rollback, feature flag, config change) while the proper fix is made, and leaves that decision to the developer.

Throughout: read the code before making a claim about it, run real commands and quote their real output, change only what the bug requires, and never push, merge or open a pull request without explicit approval.

## Steps

Work through these steps in order. Do not skip a gate.

1. reproduce (discover)
2. root-cause (discover)
3. regression-test (verify)
4. fix (build)
5. pull-request (ship)

### Step 1: Reproduce

Turn the report into a reproduction you can run on demand.

1. Restate the bug as observed versus expected behaviour. If a missing fact (version, input data, account state, configuration) blocks reproduction and the code, logs and history cannot supply it, ask for it in one message and stop.
2. Find the code path involved, from the entry point (route, command, handler, job) to the functions the symptoms point to. Cite file paths.
3. Reproduce it in the smallest form you can: a failing test or command is best, numbered manual steps are the fallback. Remove every condition that is not needed and list the ones that are.
4. Run it at least twice. If it fails only sometimes, say how often.
5. If you cannot reproduce it, do not guess a fix: report what you tried, the setup differences that could matter, and what information or instrumentation would most likely make it reproducible.
6. For high or critical severity, say who is affected now and whether a mitigation (rollback, flag, config change) would stop the harm meanwhile. Recommend it; do not apply it.

Report: the bug in one sentence, the exact reproduction with its quoted output, the required conditions, reproduced (yes, intermittent with rate, or no), and the mitigation if relevant.

Stop and wait for approval.

**Gate:** stop here and wait for the user's approval before step 2 (root-cause).

### Step 2: Root cause

Find why it happens, not just where it shows up.

1. List at most three hypotheses, ranked by how well each explains every symptom, including which conditions are required and which are not.
2. Test them one at a time with the cheapest experiment that tells them apart: a log line or breakpoint, a changed input, `git bisect` against a known-good version, a smaller reproduction. Change one thing per experiment and record the result.
3. Follow the chain to the decision in the code, data or configuration that is wrong, and explain the path from it to the symptom.
4. Ask once more why it was possible (a missing validation, a wrong assumption about an API, an unhandled state), because that decides whether the fix is local or belongs at a boundary.
5. Search for the same pattern elsewhere and list the places. Do not fix them yet.

Report: the root cause with `path:line` references, each experiment and its result, the hypotheses ruled out, why it was possible, the same pattern elsewhere, and one to three fix options with scope and risk, recommending one.

Stop and wait for approval of the cause and the fix option.

**Gate:** stop here and wait for the user's approval before step 3 (regression-test).

### Step 3: Regression test

Write the test that proves the bug, before changing the code under test.

1. Pick the cheapest level that reaches the root cause: unit if the faulty decision is in one function, integration if it lives between components or in the database, end-to-end only if nothing smaller can reach it.
2. Follow the project's test conventions; read a neighbouring test first.
3. Name the test after the behaviour, not the ticket, and assert on the outcome the user cares about with a message that explains the failure.
4. Make it deterministic: fixed clocks, seeds and data, no sleeps. For an intermittent bug, force the bad timing instead of hoping to hit it.
5. Run it against the unfixed code and confirm it fails on the bug's assertion, not on setup.

Report: the test's path, name and code, and the quoted failure with why it is the bug.

Stop and wait for approval before changing the code under test.

**Gate:** stop here and wait for the user's approval before step 4 (fix).

### Step 4: Fix

Make the smallest change that fixes the root cause.

1. Implement the approved option at the root cause. No special-casing the test's inputs, no catch-and-ignore, no retries that hide the failure, no unrelated refactors or formatting.
2. Run the regression test and confirm it passes. Then run the module's tests (the full suite if it is reasonably fast), the type check and the linter. If something unrelated was already failing, show that it fails on the original code too.
3. If callers may rely on changed behaviour (an error type, a return value, a default), list them and say whether they need updating.
4. Remove any temporary instrumentation from step 2.

Report: the diff with a line per hunk, every check with its real result, behaviour changes for callers, and anything noticed but not changed.

Stop and wait for approval before preparing the pull request.

**Gate:** stop here and wait for the user's approval before step 5 (pull-request).

### Step 5: Verify and prepare the pull request

1. Run the original reproduction from step 1 again and confirm the bug is gone. Quote the output. If the app can be run locally, check the behaviour once as the reporter would.
2. On a branch named after the behaviour (for example `fix/expired-discount-accepted`), commit the test and the fix with a message that says what was wrong and why, following the project's commit conventions.
3. Write the pull request description: the problem as the user saw it with the report's link or id; the root cause in two or three sentences; the fix and why it belongs there; the regression test and proof it failed before; risk and rollout notes (caller changes, what to watch, any mitigation to remove); and follow-ups (the same pattern elsewhere, things noticed but not changed).
4. Show the branch, commit and description. Push and open the pull request only if the developer says so; otherwise give them the commands.
````

---

<a id="debug-network-request"></a>

## Debug a failing network request

`debug-network-request` · prompt · Debugging · https://hermes-ide.com/prompts/debug-network-request

Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.

````markdown
<context>
A failing request can break at any layer between the client and the handler: name resolution, the TCP connection, TLS, a proxy or corporate gateway, the browser's CORS and mixed-content rules, authentication, timeouts at any hop, or the server rejecting the payload. Error messages from clients often hide which layer failed ("Network Error", "Failed to fetch", "socket hang up"), and people fix the wrong layer: adding CORS headers to a request that actually failed on TLS, or retrying a 401. Walking the layers in order, with one command that proves or rules out each, finds the cause quickly.
</context>

<task>
Diagnose this failing request:

<error>
[ERROR]
</error>

1. Read the error precisely and decide which layer it points to: an HTTP status means the server (or a proxy in front of it) answered, so connection, DNS and TLS worked; a browser CORS message means the request may have succeeded server-side and the browser blocked the response; connection refused, reset or timed out, certificate and name-resolution errors point lower. Say what the error rules out as well as what it suggests.
2. Walk the layers from the one most likely at fault, and for each give one command or check, what output to expect if the layer is fine, and what output means it is the problem:
   - DNS: `dig` or `nslookup` from the same machine or container, split-horizon DNS, `/etc/hosts`, stale caches.
   - Connection: `curl -v` or `nc -vz host port`; firewalls, security groups, network policies, wrong port, IPv6 versus IPv4.
   - TLS: `openssl s_client -connect host:443 -servername host`; expired or incomplete certificate chain, SNI, hostname mismatch, client trust store (corporate proxies that re-sign traffic, runtimes with their own CA bundle).
   - Proxies and gateways: `HTTP_PROXY`, `HTTPS_PROXY` and `NO_PROXY`, API gateways, header and body size limits, redirects that change the method or drop headers.
   - Browser rules: the preflight `OPTIONS` request and its `Access-Control-Allow-*` response headers, credentials with a wildcard origin, mixed content, cookies' `SameSite` and `Secure` attributes. CORS is fixed on the server, never in the client.
   - Authentication: missing or expired token, wrong audience or scope, clock skew, header stripped by a redirect or proxy, 401 versus 403 meaning.
   - Timeouts: which hop timed out (client, load balancer idle timeout, gateway, upstream), and the configured values at each.
   - Payload: content type versus body format, encoding, size, schema validation errors in a 400 or 422 body.
3. Reproduce outside the client with `curl` when possible, copying the browser request ("Copy as cURL") or translating the client's request, so client-library behaviour is separated from the server's. Say what differences between the two would be meaningful.
4. When the cause is found, give the fix at the right layer and how to confirm it.

Ask for the specific output of the next command when you need it, one or two commands at a time, rather than requesting everything up front. If the error suggests several layers equally, start with the cheapest check.
</task>

<constraints>
- Never recommend disabling TLS verification, setting a wildcard CORS origin with credentials, or turning off browser security as a fix. If used to narrow down a cause locally, label it a temporary diagnostic and never for production.
- Tell the user to remove tokens, cookies and API keys from anything they paste; use placeholders in commands.
- Give commands for the platform where the request runs (inside the container or pod if that is where it fails).
- Lead with the answer. Add reasoning only where it changes what the reader will do.
- No preamble, no restating the request and no closing summary on a short answer.
</constraints>

<output_format>
## Most likely layer
One sentence with the reason.

## What the error tells us
Two or three bullets: what it rules in and what it rules out.

## Next commands
Numbered. Each: the command in a code block, the healthy output, and the output that confirms the problem.

## Fix
Only once the cause is clear: the change, at which layer, and how to confirm it. Otherwise "Pending the output above."

## If that was not it
The next layer to check and why.
</output_format>
````

---

<a id="debug-mobile-crash"></a>

## Debug a mobile app crash

`debug-mobile-crash` · prompt · Debugging · https://hermes-ide.com/prompts/debug-mobile-crash

Debugs a mobile app crash from a symbolicated report, reading the crashed thread and frames to find the likely cause, a reproduction and a fix. Use when a crash shows up in the crash reporter.

````markdown
<context>
Mobile crash reports carry more signal than they first appear to: the exception type and signal (EXC_BAD_ACCESS with SIGSEGV, EXC_BREAKPOINT from a Swift runtime trap such as a force unwrap or array index out of range, a watchdog termination code, an uncaught Java or Kotlin exception, a native SIGABRT, an ANR's main-thread state), the crashed thread versus the main thread, the first frame in the app's own code, and the device, OS and app version spread. Cross-platform frameworks add layers: a React Native or Flutter crash may surface as a native frame, a JavaScript or Dart error, or a bridge or platform-channel call. Unsymbolicated addresses are not readable; the fix then is symbolication, not guessing.
</context>

<task>
Debug this ios crash:
<crash_report>
[CRASH_REPORT]
</crash_report>

1. Check the report matches ios. If it clearly comes from another platform (Java frames under "ios", for example), follow the report and say so. Then check it is symbolicated. If the app's frames are raw addresses, stop analysing them and explain how to symbolicate for ios (dSYMs for iOS, the R8 or ProGuard mapping file and native debug symbols for Android, Hermes or JavaScript source maps for React Native, `--split-debug-info` symbols for Flutter).
2. Read the report: exception type and signal or exception class, the reason message, the crashed thread and whether it is the main thread, the top frames, and the first frame in app code. Note what other threads were doing if a deadlock, watchdog or ANR is involved.
3. Name the crash class and what typically causes it on ios: force unwrap or out-of-range access, use after free or a dangling delegate, UI work off the main thread, main-thread blocking (watchdog or ANR), out-of-memory, a fragment or activity lifecycle state error, a null from a platform API, a JavaScript exception thrown across the bridge, a Dart null-check or platform-channel error.
4. If you can read the source, open the files in the app frames and identify the line and the conditions that lead there. Give the most likely cause with your confidence, and the next most likely if the evidence fits more than one.
5. Propose a reproduction: device or OS, steps, and conditions (slow network, backgrounding during a request, rotation, low memory, a specific locale or account state). Use breadcrumbs and the version spread to narrow it.
6. Propose the fix at the cause (not a try or catch that hides it), plus a regression test or a debug assertion where feasible.
7. Say how to verify after release: crash-free rate for the affected version, the specific crash group, and a staged rollout.
</task>

<constraints>
- Base every claim on the frames and fields in the report or on code you read. Mark anything else as a hypothesis.
- Do not suggest catching and ignoring the exception as the fix. A guard is acceptable only when the invalid state is genuinely expected, and say why it is.
- If the crash is in a third-party SDK frame, say so, check whether app code calls into it incorrectly, and suggest checking the SDK's known issues and version.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Summary
Two sentences: what crashes, the likely cause and confidence.
## Reading the report
Bullets: exception, thread, key frames with the first app frame.
## Likely cause
Explanation, with the alternative if any.
## Reproduction
Numbered steps and conditions.
## Fix
Code diff or snippet with file path, and why it addresses the cause.
## Verification
Test to add and post-release checks.
## Missing information
What would raise confidence, or "None".
</output_format>
````

---

<a id="debug-production-only-bug"></a>

## Debug a production-only bug

`debug-production-only-bug` · prompt · Debugging · https://hermes-ide.com/prompts/debug-production-only-bug

Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.

````markdown
<context>
When a bug appears only in production, the code is usually the same and something around it is not: a configuration value, a dependency version resolved differently, the data (size, shape, encoding, old records written by an earlier version), the traffic (concurrency, retries, request size), the infrastructure (proxies, load balancers, timeouts, memory limits, multiple instances), or time (time zones, clock skew, scheduled jobs, certificates or tokens expiring). Guessing and redeploying wastes days. The faster path is to list what differs, rank which difference can explain every symptom, and confirm with instrumentation that is safe to run against real users.
</context>

<task>
Debug this production-only problem:
[SYMPTOMS]

1. Extract the facts from the symptoms and logs: what fails, for whom (all users, some tenants, some regions, some instances), how often, since when, and what changed around that time (deploys, config changes, traffic growth, dependency updates, data migrations). Note patterns: specific instances, times of day, request sizes, user cohorts.
2. Diff production against the environment where it works, across these dimensions, and mark each as known-same, known-different or unknown:
   - Build and versions: commit, build flags, resolved dependency versions (lockfile honoured?), runtime and OS image, CPU architecture.
   - Configuration: environment variables, secrets, feature flags, defaults that differ when a variable is missing.
   - Data: volume, records written by older versions, nulls and unusual encodings, collation and time-zone settings, cache contents.
   - Traffic: concurrency, request sizes, retries, long-lived connections, bots.
   - Infrastructure: multiple instances (local state, sticky sessions), proxies and load balancers (header size, body size, idle timeouts), network policies, DNS, memory and CPU limits, file-system permissions and read-only volumes.
   - Time: time zones, clock skew between hosts, daylight saving, scheduled jobs, expiring certificates or tokens.
   - Dependencies: third-party API behaviour in production versus sandbox, rate limits, regional endpoints.
3. Form at most four hypotheses. For each, say which symptoms it explains and which it does not; drop hypotheses that contradict the evidence.
4. For each remaining hypothesis, design the cheapest confirming check, in order of safety: read-only queries and comparisons first (compare configs, query the data, read existing logs and metrics), then reproduction with production-like conditions in a non-production environment (production data snapshot with personal data masked, same versions, load), and only then targeted production instrumentation: extra log fields or spans behind a flag, sampled, for a limited time, on a subset of traffic, with no personal data or secrets logged and a plan to remove it.
5. Give the likely fix for the leading hypothesis and how to verify it in production after release (which metric or log should change).

If the symptoms are too vague to form any hypothesis, ask the three questions whose answers would narrow it most, and stop.
</task>

<constraints>
- Do not suggest attaching a debugger to production, enabling verbose logging globally, or experimenting on production data. Production instrumentation must be scoped, sampled, time-boxed and free of personal data.
- Each hypothesis must account for why it does not happen in the working environment.
- Never ask for secrets or credentials; ask for whether a value is set or how it differs.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## What the evidence says
Bullets of facts, each with its source (symptom report, log line, metric).

## Differences that matter
Table: dimension | production | working environment | known-same, known-different or unknown.

## Hypotheses
Numbered, most likely first. Each: the cause, the symptoms it explains, the ones it does not, and why the working environment is unaffected.

## Confirm safely
Per hypothesis, the checks in order with exactly what to run or look at and what result confirms or rules it out.

## Likely fix
The fix for the leading hypothesis and the production signal that proves it worked.

## Missing information
What to collect next, most useful first.
</output_format>
````

---

<a id="debug-race-condition"></a>

## Debug a race condition

`debug-race-condition` · prompt · Debugging · https://hermes-ide.com/prompts/debug-race-condition

Diagnoses intermittent concurrency bugs by mapping shared state and the interleavings that break it, adds targeted instrumentation and proposes a fix. Use for bugs seen only under load.

````markdown
<context>
Race conditions are bugs in ordering: two or more units of execution touch the same state, and some interleaving of their steps breaks an invariant. They hide from debuggers and print statements because observing them changes the timing. The reliable way in is to reason from the shared state and the possible interleavings, form specific hypotheses, then make the bad interleaving more likely on purpose and prove it with evidence. Sleeps, retries and "add a lock somewhere" usually move the bug rather than remove it.
</context>

<task>
Diagnose this concurrency bug.

Code:
[CODE]

Symptoms:
[SYMPTOMS]


If the runtime or the concurrency model is not clear from the code, ask before going further, because the answer changes which interleavings are possible.

1. Map the concurrency: list each unit that runs concurrently (threads, goroutines, async tasks, workers, processes, app instances, cron jobs) and each piece of shared state (in-memory fields, caches, globals, files, database rows, queues, external resources). For each piece, list every read and write with its location and the synchronisation that protects it, if any.
2. Name the invariant that the symptom shows is broken (for example, "an order is charged at most once").
3. Enumerate candidate interleavings that break it. Check at least: check-then-act and read-modify-write without atomicity; lost updates in the database under the actual isolation level; publication without a happens-before edge (unsafe lazy init, non-volatile flags); iterating a collection while it is modified; await points that split a critical section in single-threaded async code; lock ordering that can deadlock; time-of-check to time-of-use on files or external state; duplicate delivery from retries or at-least-once queues. Write each candidate as a step-by-step timeline of A and B.
4. Rank candidates by how well they explain every symptom (frequency, load dependence, the exact wrong value). Drop those that contradict the evidence.
5. Propose instrumentation that can confirm or rule out the top candidates without hiding the bug: log lines with unit id, monotonic timestamp and a sequence or version number at each access; the runtime's race detector or concurrency checker if one exists for this runtime; a stress test that runs the operation concurrently many times, with injected delays or yields at the suspected gap to widen the window.
6. Propose the fix that removes the bad interleaving at its root, preferring in order: removing the sharing, making the operation atomic (a single atomic op, a conditional update, a unique constraint, a transaction at the right isolation level, optimistic locking with a version), then a lock with a documented scope and order. Make operations idempotent where duplicates are possible.
7. Define how to verify: the stress test fails before the fix at a measured rate and passes after many runs.
</task>

<constraints>
- Never propose sleeps, retries or longer timeouts as the fix.
- Do not claim a root cause is confirmed until the evidence from step 5 confirms it; until then, call it the leading hypothesis.
- Keep the fix as small as the root cause allows, and state what it costs (contention, throughput, latency).
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Shared state
Table: State | Readers and writers (location) | Protection.
## Candidate interleavings
Ranked. For each: the broken invariant, a two-column timeline (A | B), and how well it explains the symptoms.
## Instrumentation
What to add or run, and the result that would confirm or rule out each top candidate.
## Fix
The diff for the leading candidate, and why it removes the interleaving.
## Verification
The stress or race-detector test, how many runs, and the before and after failure rates to expect.
</output_format>
````

---

<a id="debugger"></a>

## Debugger

`debugger` · persona · Debugging · https://hermes-ide.com/prompts/debugger

Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.

````markdown
From now on, work as this persona: Debugger.

You are a debugger. You treat every bug as a question about the difference between what the code assumes and what actually happens, and you answer it with experiments, not intuition.

How you work:
- You reproduce first. A failure you can trigger on demand, ideally with one command or one failing test, comes before any theory.
- You keep observations and assumptions apart, and you write both down as you go.
- You hold several hypotheses at once and pick the experiment that best separates them, usually the cheapest one: a log line, an assertion, a changed input, a bisect over commits or data.
- You change one thing at a time and predict the result before you run it. A surprise means your model of the system is wrong, and that is useful.
- You stop when you can predict the failure, not when you have a plausible story.

What you flag:
- Symptom fixes: swallowed exceptions, added retries or sleeps, null checks where the null should never arrive, special cases for one input.
- Assumptions nobody checked: time zones, encodings, ordering, caching, environment differences between machines.
- Missing information: when a report or log cannot settle the question, you say exactly what would.
- Errors in the code that reports errors: lost stack traces, rethrown exceptions without the cause, misleading messages.

Your habits:
- You fix the cause with the smallest change, remove the instrumentation you added, and leave a test that fails without the fix.
- You show your evidence: the command, the output, the before and after.
- You say "I don't know yet" when you don't, together with the next experiment.
- You never touch someone's uncommitted work without asking.
````

---

<a id="explain-stack-trace"></a>

## Explain a stack trace

`explain-stack-trace` · prompt · Debugging · https://hermes-ide.com/prompts/explain-stack-trace

Explains an error and its stack trace in plain words, finds the frame that matters, and ranks the likely causes with the next checks to run. Use when an exception or crash is hard to read.

````markdown
<context>
Stack traces are long, and most of their frames belong to frameworks and libraries. The useful information is usually three things: the real exception (often the innermost one in a chain), the first frame in the project's own code, and the value that was wrong when it got there. Each runtime prints these differently.
</context>

<task>
Explain this error:
[TRACE]
1. Identify the language or runtime from the trace format, and read the trace in that runtime's order:
   - Python prints the most recent call last, so the failing line is at the bottom.
   - Java, Kotlin and C# put the outermost exception first; the root is the last "Caused by" or inner exception.
   - JavaScript and TypeScript traces may be cut at async boundaries and may point to compiled files; say when a source map is needed.
   - Go panics list each goroutine; the panicking goroutine comes first. Rust panics need `RUST_BACKTRACE=1` for a full trace.
2. Find the root exception and its message. Say what it means in one plain sentence.
3. Find the first frame in the project's own code, as opposed to the standard library, a framework or a dependency. If the project's code is available, read that line and the lines that feed it.
4. Reason backwards from that line: which value or state must have been wrong for this error to happen, and where could it have come from?
5. Rank the likely causes and give the cheapest check that confirms or rules out each one.
</task>

<constraints>
- Do not guess at code you have not seen. If the project's code is not available, base the explanation on the trace alone and say so.
- Quote frames exactly as they appear in the trace. Never invent file names, line numbers or function names.
- Ignore framework and library frames unless the error originates inside one. If it does, say whether the likely fault is still the caller's input.
- If the trace is truncated or minified so that the cause cannot be found, say what is missing and how to get it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
- Lead with the answer. Add reasoning only where it changes what the reader will do.
- No preamble, no restating the request and no closing summary on a short answer.
</constraints>

<output_format>
## What happened
One or two plain sentences: the root exception and what it means.
## Where
The frame that matters, quoted from the trace, and why that frame.
## Likely causes
Numbered, most likely first. Each cause with the evidence for it.
## Next checks
Bullets: one concrete check per cause (a value to print, a line to read, a command to run).
</output_format>
````

---

<a id="find-root-cause"></a>

## Find the root cause of a bug

`find-root-cause` · prompt · Debugging · https://hermes-ide.com/prompts/find-root-cause

Reproduces a bug, tests ranked hypotheses with experiments, and fixes the root cause instead of the symptom. Use when something is broken and the reason is not obvious.

````markdown
<context>
A fix that targets the symptom usually moves the bug instead of removing it: a null check where the null should never arrive, a retry around a race, a catch that hides the error. The root cause is the earliest point where the program's actual state diverges from what the code assumes. Debugging is finding that point with experiments, not guessing at it.
</context>

<task>
Find and fix the root cause of: [SYMPTOM]
1. **Reproduce.** Find the shortest reliable way to trigger the symptom, ideally a single command or a failing test. Record how often it fails. If you cannot reproduce it, say what you tried and what information would let you, then stop and ask.
2. **Collect facts.** Read the code on the failing path. Separate what you observed (outputs, logs, values) from what you assume.
3. **Hypothesise.** List two to five candidate causes. For each, state what you would expect to see if it were true and if it were false.
4. **Experiment.** Run the cheapest experiment that best separates the hypotheses: add a log or assertion, inspect a value, change one input, bisect the code path, the input data or the commit history. Change one thing at a time and record each result.
5. **Confirm.** You have the root cause when you can predict the failure, for example "with input X it fails; with Y it passes", and the prediction holds.
6. **Fix at the cause**, as the smallest correct change. Remove the temporary logs and assertions you added.
7. **Verify.** Run the reproduction again and the surrounding tests. Add a test that fails without the fix when the project has tests.
</task>

<constraints>
- Do not change code to "see if it helps" without a hypothesis that predicts the result.
- Do not stop at the first plausible explanation. Confirm it with an experiment whose result you predicted.
- Never fix the symptom by swallowing errors, adding retries or sleeps, or special-casing the failing input. If a symptom-level mitigation is needed urgently, label it as such and still name the root cause.
- If the cause is outside the code (configuration, data, environment, a dependency), say so and stop at a recommendation.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Reproduction
The command or steps, and the failure rate observed.
## Hypotheses
A table: Hypothesis | Experiment | Result | Verdict (confirmed, ruled out, open).
## Root cause
One paragraph: where the state first goes wrong (`path:line`), why, and how that produces the symptom.
## Fix
The diff, then one sentence on why it removes the cause.
## Verification
The commands you ran after the fix and their results, including the new test.
</output_format>
````

---

<a id="triage-failing-ci"></a>

## Triage a failing CI build

`triage-failing-ci` · prompt · Debugging · https://hermes-ide.com/prompts/triage-failing-ci

Finds the first real error in a failing CI log, classifies the failure as caused by the change, flaky, environment drift or already broken, and names the next action. Use when a pipeline turns red.

````markdown
<context>
A red build is a question with a few common answers: the change broke something, a test is flaky, the environment drifted (a new dependency release, a new runner image, an expired credential, a rate limit), or the base branch was already broken. The answer decides who acts and how. CI logs bury the first real error under cascading failures and noisy setup output.
</context>

<task>
Triage this CI failure:
[CI_LOG]
1. Find the first real error: the earliest failure that the later ones follow from. Skip warnings, deprecation notices and failures that only happen because an earlier step failed.
2. Classify the failure:
   - **change**: the error is in code, tests or config the change touched, or plainly follows from it.
   - **flaky**: timing, ordering or network-dependent failure, unrelated to the change. Look for timeouts, connection resets, port conflicts and tests that touch time or randomness.
   - **environment**: dependency versions resolved differently than before, a runner or image update, missing secrets, quota or rate limits, full disks.
   - **pre-existing**: the same failure is on the base branch. Check the base branch's recent runs or history if you can.
3. Give the evidence for the classification and what would change your mind.
4. Name the next action and who should take it: fix the code (with the likely location), rerun with a reason, pin a dependency, or report an infrastructure issue.
</task>

<constraints>
- Quote the first real error exactly, with its step name and line in the log if available.
- Recommend a rerun only for **flaky** or transient **environment** failures, and say why. Never recommend rerunning a deterministic failure.
- Do not recommend disabling or skipping a test unless the test itself is proven to be broken, and then say how to track re-enabling it.
- If the log is truncated before the error, say so and say which part of the log you need.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Classification
`change`, `flaky`, `environment` or `pre-existing`, with confidence (high, medium, low).
## First real error
The quoted error, its job and step.
## Evidence
Bullets supporting the classification, and one line on what would change it.
## Next action
One or two concrete steps, with the likely file or setting to look at.
</output_format>
````

---

<a id="reproduce-bug-report"></a>

## Turn a bug report into a minimal reproduction

`reproduce-bug-report` · prompt · Debugging · https://hermes-ide.com/prompts/reproduce-bug-report

Turns a vague bug report into a minimal, reliable reproduction, preferably a failing test, and states the exact conditions needed. Use before fixing a reported bug or when triaging issues.

````markdown
<context>
A bug that cannot be reproduced cannot be fixed with confidence. Reports mix what the user saw with what they think caused it, and they leave out the conditions that matter. A minimal reproduction strips everything that is not needed to trigger the failure, which often points straight at the cause.
</context>

<task>
Reproduce this report:
[REPORT]
1. Separate the report into observations (what the user saw) and interpretations (what they think caused it). Work from the observations.
2. Write down the expected and the actual behaviour in one line each. If the report does not make expected behaviour clear, say so.
3. Reproduce it in the codebase, starting at the closest level you can: a unit or integration test first, then a script or command, and manual steps only as a last resort.
4. Minimise: remove inputs, steps and configuration one at a time while the failure still happens. Then vary the conditions that seem to matter (data shape, version, platform, timing, configuration) to find which ones are required.
5. Leave the reproduction in place as a failing test, marked so it is easy to find, or as exact steps if a test is not possible.
</task>

<constraints>
- Do not fix the bug. This task ends at a reliable reproduction.
- If you cannot reproduce it, do not pretend you did. List the attempts and the conditions you tried, and write the questions for the reporter that would unblock you.
- Keep the reproduction free of real user data. Use synthetic values with the same shape.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Status
`reproduced`, `partly reproduced` or `not reproduced`, and the failure rate when it is intermittent.
## Reproduction
The failing test (path and code) or the exact steps and command, and its output.
## Conditions
Bullets: what must be true for the failure to happen, and what turned out not to matter.
## Expected and actual
Two lines.
## Unknowns
Questions for the reporter, or "None".
</output_format>
````
