hermes

Debug a production-only bug

Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.

context

When a bug appears only in production, the code is usually the same and something around it is not: a configuration value, a dependency version resolved differently, the data (size, shape, encoding, old records written by an earlier version), the traffic (concurrency, retries, request size), the infrastructure (proxies, load balancers, timeouts, memory limits, multiple instances), or time (time zones, clock skew, scheduled jobs, certificates or tokens expiring). Guessing and redeploying wastes days. The faster path is to list what differs, rank which difference can explain every symptom, and confirm with instrumentation that is safe to run against real users.

task

Debug this production-only problem: Only if [ENVIRONMENT_DIFFERENCES] is given:

Known differences: Only if [LOGS] is given:

logs

  1. Extract the facts from the symptoms and logs: what fails, for whom (all users, some tenants, some regions, some instances), how often, since when, and what changed around that time (deploys, config changes, traffic growth, dependency updates, data migrations). Note patterns: specific instances, times of day, request sizes, user cohorts.
  2. Diff production against the environment where it works, across these dimensions, and mark each as known-same, known-different or unknown:
  • Build and versions: commit, build flags, resolved dependency versions (lockfile honoured?), runtime and OS image, CPU architecture.
  • Configuration: environment variables, secrets, feature flags, defaults that differ when a variable is missing.
  • Data: volume, records written by older versions, nulls and unusual encodings, collation and time-zone settings, cache contents.
  • Traffic: concurrency, request sizes, retries, long-lived connections, bots.
  • Infrastructure: multiple instances (local state, sticky sessions), proxies and load balancers (header size, body size, idle timeouts), network policies, DNS, memory and CPU limits, file-system permissions and read-only volumes.
  • Time: time zones, clock skew between hosts, daylight saving, scheduled jobs, expiring certificates or tokens.
  • Dependencies: third-party API behaviour in production versus sandbox, rate limits, regional endpoints.
  1. Form at most four hypotheses. For each, say which symptoms it explains and which it does not; drop hypotheses that contradict the evidence.
  2. For each remaining hypothesis, design the cheapest confirming check, in order of safety: read-only queries and comparisons first (compare configs, query the data, read existing logs and metrics), then reproduction with production-like conditions in a non-production environment (production data snapshot with personal data masked, same versions, load), and only then targeted production instrumentation: extra log fields or spans behind a flag, sampled, for a limited time, on a subset of traffic, with no personal data or secrets logged and a plan to remove it.
  3. Give the likely fix for the leading hypothesis and how to verify it in production after release (which metric or log should change).

If the symptoms are too vague to form any hypothesis, ask the three questions whose answers would narrow it most, and stop.

constraints
  • Do not suggest attaching a debugger to production, enabling verbose logging globally, or experimenting on production data. Production instrumentation must be scoped, sampled, time-boxed and free of personal data.
  • Each hypothesis must account for why it does not happen in the working environment.
  • Never ask for secrets or credentials; ask for whether a value is set or how it differs.
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

What the evidence says

Bullets of facts, each with its source (symptom report, log line, metric).

Differences that matter

Table: dimension | production | working environment | known-same, known-different or unknown.

Hypotheses

Numbered, most likely first. Each: the cause, the symptoms it explains, the ones it does not, and why the working environment is unaffected.

Confirm safely

Per hypothesis, the checks in order with exactly what to run or look at and what result confirms or rules it out.

Likely fix

The fix for the leading hypothesis and the production signal that proves it worked.

Missing information

What to collect next, most useful first.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Debugging
level
Intermediate
made for
Backend engineer, Software engineer, Site reliability engineer, DevOps / platform engineer
needs
repo-read
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install debug-production-only-bug --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill debug-production-only-bug -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Debugging
PromptIncident and operations

Instrument a service for observability

Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.

instrument-service-observability
PromptDebugging

Find the root cause of a bug

Reproduces a bug, tests ranked hypotheses with experiments, and fixes the root cause instead of the symptom. Use when something is broken and the reason is not obvious.

find-root-cause
PromptDebugging

Debug a race condition

Diagnoses intermittent concurrency bugs by mapping shared state and the interleavings that break it, adds targeted instrumentation and proposes a fix. Use for bugs seen only under load.

debug-race-condition
PromptIncident and operations

Triage a production alert

Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.

triage-production-alert
PersonaDebugging

Debugger

Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.

debugger
PromptDebugging

Debug a failing network request

Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.

debug-network-request