Debug a production-only bug
Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.
When a bug appears only in production, the code is usually the same and something around it is not: a configuration value, a dependency version resolved differently, the data (size, shape, encoding, old records written by an earlier version), the traffic (concurrency, retries, request size), the infrastructure (proxies, load balancers, timeouts, memory limits, multiple instances), or time (time zones, clock skew, scheduled jobs, certificates or tokens expiring). Guessing and redeploying wastes days. The faster path is to list what differs, rank which difference can explain every symptom, and confirm with instrumentation that is safe to run against real users.
Debug this production-only problem: Only if [ENVIRONMENT_DIFFERENCES] is given:
Known differences: Only if [LOGS] is given:
- Extract the facts from the symptoms and logs: what fails, for whom (all users, some tenants, some regions, some instances), how often, since when, and what changed around that time (deploys, config changes, traffic growth, dependency updates, data migrations). Note patterns: specific instances, times of day, request sizes, user cohorts.
- Diff production against the environment where it works, across these dimensions, and mark each as known-same, known-different or unknown:
- Build and versions: commit, build flags, resolved dependency versions (lockfile honoured?), runtime and OS image, CPU architecture.
- Configuration: environment variables, secrets, feature flags, defaults that differ when a variable is missing.
- Data: volume, records written by older versions, nulls and unusual encodings, collation and time-zone settings, cache contents.
- Traffic: concurrency, request sizes, retries, long-lived connections, bots.
- Infrastructure: multiple instances (local state, sticky sessions), proxies and load balancers (header size, body size, idle timeouts), network policies, DNS, memory and CPU limits, file-system permissions and read-only volumes.
- Time: time zones, clock skew between hosts, daylight saving, scheduled jobs, expiring certificates or tokens.
- Dependencies: third-party API behaviour in production versus sandbox, rate limits, regional endpoints.
- Form at most four hypotheses. For each, say which symptoms it explains and which it does not; drop hypotheses that contradict the evidence.
- For each remaining hypothesis, design the cheapest confirming check, in order of safety: read-only queries and comparisons first (compare configs, query the data, read existing logs and metrics), then reproduction with production-like conditions in a non-production environment (production data snapshot with personal data masked, same versions, load), and only then targeted production instrumentation: extra log fields or spans behind a flag, sampled, for a limited time, on a subset of traffic, with no personal data or secrets logged and a plan to remove it.
- Give the likely fix for the leading hypothesis and how to verify it in production after release (which metric or log should change).
If the symptoms are too vague to form any hypothesis, ask the three questions whose answers would narrow it most, and stop.
- Do not suggest attaching a debugger to production, enabling verbose logging globally, or experimenting on production data. Production instrumentation must be scoped, sampled, time-boxed and free of personal data.
- Each hypothesis must account for why it does not happen in the working environment.
- Never ask for secrets or credentials; ask for whether a value is set or how it differs.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
What the evidence says
Bullets of facts, each with its source (symptom report, log line, metric).
Differences that matter
Table: dimension | production | working environment | known-same, known-different or unknown.
Hypotheses
Numbered, most likely first. Each: the cause, the symptoms it explains, the ones it does not, and why the working environment is unaffected.
Confirm safely
Per hypothesis, the checks in order with exactly what to run or look at and what result confirms or rules it out.
Likely fix
The fix for the leading hypothesis and the production signal that proves it worked.
Missing information
What to collect next, most useful first.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Debugging
- level
- Intermediate
- made for
- Backend engineer, Software engineer, Site reliability engineer, DevOps / platform engineer
- needs
- repo-read
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install debug-production-only-bug --target claude-codenpx skills add hermes-hq/hodios-dist --skill debug-production-only-bug -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of DebuggingInstrument a service for observability
Plans and adds logs, metrics and traces using OpenTelemetry conventions, golden signals, useful log fields, cardinality limits and first dashboards. Use when a service is hard to debug in production.
instrument-service-observabilityFind the root cause of a bug
Reproduces a bug, tests ranked hypotheses with experiments, and fixes the root cause instead of the symptom. Use when something is broken and the reason is not obvious.
find-root-causeDebug a race condition
Diagnoses intermittent concurrency bugs by mapping shared state and the interleavings that break it, adds targeted instrumentation and proposes a fix. Use for bugs seen only under load.
debug-race-conditionTriage a production alert
Turns a firing production alert into a severity call, the safest mitigation to try first, ranked hypotheses and the next checks. Use in the first minutes of an incident or page.
triage-production-alertDebugger
Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.
debuggerDebug a failing network request
Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.
debug-network-request