Triage a failing CI build
Finds the first real error in a failing CI log, classifies the failure as caused by the change, flaky, environment drift or already broken, and names the next action. Use when a pipeline turns red.
A red build is a question with a few common answers: the change broke something, a test is flaky, the environment drifted (a new dependency release, a new runner image, an expired credential, a rate limit), or the base branch was already broken. The answer decides who acts and how. CI logs bury the first real error under cascading failures and noisy setup output.
Triage this CI failure: Only if [CHANGE] is given: Triggering change:
- Find the first real error: the earliest failure that the later ones follow from. Skip warnings, deprecation notices and failures that only happen because an earlier step failed.
- Classify the failure:
- change: the error is in code, tests or config the change touched, or plainly follows from it.
- flaky: timing, ordering or network-dependent failure, unrelated to the change. Look for timeouts, connection resets, port conflicts and tests that touch time or randomness.
- environment: dependency versions resolved differently than before, a runner or image update, missing secrets, quota or rate limits, full disks.
- pre-existing: the same failure is on the base branch. Check the base branch's recent runs or history if you can.
- Give the evidence for the classification and what would change your mind.
- Name the next action and who should take it: fix the code (with the likely location), rerun with a reason, pin a dependency, or report an infrastructure issue.
- Quote the first real error exactly, with its step name and line in the log if available.
- Recommend a rerun only for flaky or transient environment failures, and say why. Never recommend rerunning a deterministic failure.
- Do not recommend disabling or skipping a test unless the test itself is proven to be broken, and then say how to track re-enabling it.
- If the log is truncated before the error, say so and say which part of the log you need.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Classification
change, flaky, environment or pre-existing, with confidence (high, medium, low).
First real error
The quoted error, its job and step.
Evidence
Bullets supporting the classification, and one line on what would change it.
Next action
One or two concrete steps, with the likely file or setting to look at.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Debugging
- level
- Intermediate
- made for
- Software engineer, DevOps / platform engineer
- needs
- repo-read, git
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md
use in
npx @hermes-hq/hodios install triage-failing-ci --target claude-codenpx skills add hermes-hq/hodios-dist --skill triage-failing-ci -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of DebuggingDebugger
Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.
debuggerFix a flaky test
Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.
fix-flaky-testFind the root cause of a bug
Reproduces a bug, tests ranked hypotheses with experiments, and fixes the root cause instead of the symptom. Use when something is broken and the reason is not obvious.
find-root-causeDebug a failing network request
Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.
debug-network-requestDebug a production-only bug
Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.
debug-production-only-bugBisect a regression
Finds the commit or input that introduced a regression by writing an automated good/bad check first, then bisecting. Use when something that used to work is broken and the cause is unclear.
bisect-regression