hermes

Triage a failing CI build

Finds the first real error in a failing CI log, classifies the failure as caused by the change, flaky, environment drift or already broken, and names the next action. Use when a pipeline turns red.

context

A red build is a question with a few common answers: the change broke something, a test is flaky, the environment drifted (a new dependency release, a new runner image, an expired credential, a rate limit), or the base branch was already broken. The answer decides who acts and how. CI logs bury the first real error under cascading failures and noisy setup output.

task

Triage this CI failure: Only if [CHANGE] is given: Triggering change:

  1. Find the first real error: the earliest failure that the later ones follow from. Skip warnings, deprecation notices and failures that only happen because an earlier step failed.
  2. Classify the failure:
  • change: the error is in code, tests or config the change touched, or plainly follows from it.
  • flaky: timing, ordering or network-dependent failure, unrelated to the change. Look for timeouts, connection resets, port conflicts and tests that touch time or randomness.
  • environment: dependency versions resolved differently than before, a runner or image update, missing secrets, quota or rate limits, full disks.
  • pre-existing: the same failure is on the base branch. Check the base branch's recent runs or history if you can.
  1. Give the evidence for the classification and what would change your mind.
  2. Name the next action and who should take it: fix the code (with the likely location), rerun with a reason, pin a dependency, or report an infrastructure issue.
constraints
  • Quote the first real error exactly, with its step name and line in the log if available.
  • Recommend a rerun only for flaky or transient environment failures, and say why. Never recommend rerunning a deterministic failure.
  • Do not recommend disabling or skipping a test unless the test itself is proven to be broken, and then say how to track re-enabling it.
  • If the log is truncated before the error, say so and say which part of the log you need.
  • Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
  • If the information you need is not available, say what is missing and how to get it instead of inventing it.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Classification

change, flaky, environment or pre-existing, with confidence (high, medium, low).

First real error

The quoted error, its job and step.

Evidence

Bullets supporting the classification, and one line on what would change it.

Next action

One or two concrete steps, with the likely file or setting to look at.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Debugging
level
Intermediate
made for
Software engineer, DevOps / platform engineer
needs
repo-read, git
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install triage-failing-ci --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill triage-failing-ci -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Debugging
PersonaDebugging

Debugger

Debugs by reproducing first, testing one hypothesis at a time and fixing root causes, never symptoms. Use as a persona or subagent for bugs, crashes and failing builds.

debugger
PromptTesting

Fix a flaky test

Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.

fix-flaky-test
PromptDebugging

Find the root cause of a bug

Reproduces a bug, tests ranked hypotheses with experiments, and fixes the root cause instead of the symptom. Use when something is broken and the reason is not obvious.

find-root-cause
PromptDebugging

Debug a failing network request

Diagnoses a failing HTTP request layer by layer (DNS, TLS, proxy, CORS, auth, timeouts, payload) from error output and curl or browser traces, giving the next command at each step.

debug-network-request
PromptDebugging

Debug a production-only bug

Debugs a bug that happens only in production by diffing environment, config, data, traffic, versions and timing, then plans safe instrumentation to confirm the cause. Use for works-on-my-machine bugs.

debug-production-only-bug
PromptDebugging

Bisect a regression

Finds the commit or input that introduced a regression by writing an automated good/bad check first, then bisecting. Use when something that used to work is broken and the cause is unclear.

bisect-regression