Bisect a regression
Finds the commit or input that introduced a regression by writing an automated good/bad check first, then bisecting. Use when something that used to work is broken and the cause is unclear.
Bisection finds the first bad commit in log2(n) steps, but only if every step is judged correctly. Most failed bisects come from a manual or flaky check, an untestable commit marked bad, or a "good" endpoint that was never verified. So the check comes first: one script that builds what it needs, reproduces the symptom, and exits with an unambiguous code. The same idea applies when the regression is triggered by data rather than code: halve the input until the smallest failing input remains.
Find what introduced this regression:
Bad: . Only if [GOOD_REF] is given: Last known good: .
- Write the check. A script that exits 0 when the behaviour is good, 1 when it shows this specific regression, and 125 when the commit cannot be tested (build fails for an unrelated reason, missing migration). It must test the regression itself, not "any failure", and must map crashes and signals to 1 or 125 explicitly, because
git bisect runaborts on any exit code above 127. Make it deterministic: fixed seeds, clean build output, isolated temp data. If the symptom is intermittent, run it N times and call it bad if any run fails; say what N gives enough confidence for the failure rate you observed. - Confirm the endpoints. Run the check on the bad ref and the good ref and show the results. If no good ref is known, find one by testing older release tags or stepping back exponentially (bad~10, ~20, ~40…), and stop to ask if nothing older is good. If the check disagrees with the user's report on either endpoint, stop and fix the check.
- Decide what to bisect. If the regression appears with the same code and different data or configuration, bisect the input instead: split the input in halves (records, config keys, files), keep the half that still fails, and repeat until removing any single part makes it pass.
- Run the bisect:
git bisect start <bad> <good>, thengit bisect run <check>. Use--first-parentwhen the history has merges and the team wants the merge that introduced it. Note any skipped commits. - Confirm the culprit. Show the commit, read its diff, and explain the mechanism that breaks the behaviour. Where practical, revert just that commit on top of the bad ref and show the check passes.
- End with
git bisect resetand say which branch is checked out.
- Never mark a commit bad because it fails to build or fails for a different reason; that is a skip (exit 125).
- Do not modify tracked files during the bisect; keep the check script outside the repository or untracked so checkouts do not change it.
- If you cannot run commands, give the user the check script and the exact commands, and ask for the output at each decision point instead of guessing results.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Check
The script in a code block, and what each exit code means for this regression.
Good and bad endpoints
The refs and the check result on each.
Bisect
The exact commands, and the bisect log if you ran it.
Result
The first bad commit (hash, title, author date), or the minimal failing input, with how many steps it took and any skipped commits.
Culprit analysis
What in that change causes the regression, the revert check, and a suggested next step (fix forward or revert).
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Debugging
- level
- Intermediate
- made for
- Software engineer, QA / test engineer, Open-source maintainer
- needs
- repo-read, shell, git
- risk
- runs-commands
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install bisect-regression --target claude-codenpx skills add hermes-hq/hodios-dist --skill bisect-regression -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of DebuggingDebugger
Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.
debuggerDebug a failing network request
Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.
debug-network-requestDebug a production-only bug
Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.
debug-production-only-bugBugfix track
Takes a bug from report to reproduction, root cause, regression test, minimal fix and a verified pull request, stopping for approval between steps. Use for any bug worth fixing properly.
bugfix-trackDebug a mobile app crash
Debugs a mobile app crash from a symbolicated report, reading the crashed thread and frames to find the likely cause, a reproduction and a fix. Use when a crash shows up in the crash reporter.
debug-mobile-crashDebug a race condition
Diagnoses intermittent concurrency bugs by mapping shared state and the interleavings that break it, adds targeted instrumentation and proposes a fix. Use for bugs seen only under load.
debug-race-condition