hermes

Diagnose prompt failures

Diagnoses why a prompt produces bad answers from failing examples, traces each failure to a root cause, proposes targeted fixes and a quick regression test set.

context

When a prompt misbehaves, people tend to rewrite it from scratch or pile on capital-letter warnings. Both make things worse: the rewrite breaks what used to work, and the warnings make the model overcorrect elsewhere. Debugging a prompt works like debugging code: look at the failures, form hypotheses, find the root cause, make the smallest change that addresses it, and check that nothing else broke.

prompt under test

bad outputs

Only if [EXPECTED] is given:

expected

task
  1. Describe each failure precisely: what was expected, what happened, and the exact part of the output that is wrong. If no expectation is given and it is not obvious, infer it and say so.
  2. Group failures into patterns.
  3. For each pattern, test these causes against the evidence and name the most likely root cause:
  • The instruction is missing, ambiguous, or only implied.
  • Instructions conflict, or one buried late or deep is outweighed by an earlier one.
  • Examples are being copied (length, wording, labels) or do not cover the failing case.
  • Input is not delimited, so the model treats data as instructions or mixes it into the answer.
  • The output format is underspecified, or the reasoning and the final answer are mixed.
  • Missing context or knowledge, so the model fills gaps by guessing.
  • Too many jobs in one prompt.
  • Not a prompt problem: a capability limit (exact counting, long arithmetic, very long inputs), missing retrieval or tools, settings such as temperature or maximum length, or the pipeline around the model.
  1. Propose the smallest targeted fix for each root cause, show it as a before and after, and say which failures it should fix and what it might break.
  2. Give the revised prompt with all fixes applied and nothing else changed.
  3. Build a quick test set: every failing input, three to five inputs that worked before (to catch regressions), and two new edge cases, each with a pass condition that can be checked.
constraints
  • Base every diagnosis on evidence in the outputs or the prompt. If the evidence is too thin to tell causes apart, say so and propose a small experiment that would (for example, remove the examples and rerun).
  • Prefer explaining the reason behind a rule over adding emphasis.
  • Do not claim a fix works; you cannot run it. Say what result would confirm it.
  • If a cause is outside the prompt, say so plainly and recommend the right fix (a tool, retrieval, validation code, a setting) rather than more instructions.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Failure patterns

Table: Failure | Expected | Got | Pattern.

Root causes

One short paragraph per pattern with the evidence.

Fixes

Numbered. Each: Before, After, Fixes which failures, Risk.

Revised prompt

Fenced code block.

Test set

Table: Input | Why it is in the set | Pass condition.

If this does not fix it

The next hypothesis to test, and how.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Prompting and assistants
category
Prompt engineering
level
Intermediate
made for
ML / AI engineer, Software engineer, Product manager, Anyone, personal use
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install diagnose-prompt-failures --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill diagnose-prompt-failures -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the prompting plugin
claude plugin install hodios-prompting@hodios

The plugin brings every entry in this domain at once.

PersonaPrompt engineering

Prompt engineer

Prompt engineer who writes clear, testable instructions, iterates against real examples and evals, and avoids model-specific tricks. Use for designing, debugging and maintaining prompts.

prompt-engineer
PromptPrompt engineering

Improve a prompt

Diagnoses why a prompt gives weak or inconsistent results and rewrites it with clear context, task, constraints and output format while keeping its intent. Use on any prompt for any AI assistant.

improve-prompt
PromptPrompt engineering

Create few-shot examples

Builds a small set of diverse, representative few-shot examples for a task, including tricky and negative cases, balanced so the model learns the rule rather than copying surface patterns.

create-few-shot-examples
PromptPrompt engineering

Compress a prompt

Shortens a long prompt while preserving its behaviour, maps every original instruction to where it now lives, reports the real size reduction and lists test inputs to check nothing changed.

compress-prompt
PromptPrompt engineering

Adapt a prompt for a reasoning model

Rewrites a prompt for reasoning-capable models by removing step-by-step micromanagement, stating goals, constraints and success criteria, and keeping the output format exact.

adapt-prompt-for-reasoning-model
PromptPrompt engineering

Build a test set for a prompt

Builds a hand-run test set for a prompt with happy, edge and negative inputs, expected behaviour and checkable pass criteria per case, and a scoring sheet to compare prompt versions side by side.

build-prompt-test-set