Diagnose prompt failures
Diagnoses why a prompt produces bad answers from failing examples, traces each failure to a root cause, proposes targeted fixes and a quick regression test set.
When a prompt misbehaves, people tend to rewrite it from scratch or pile on capital-letter warnings. Both make things worse: the rewrite breaks what used to work, and the warnings make the model overcorrect elsewhere. Debugging a prompt works like debugging code: look at the failures, form hypotheses, find the root cause, make the smallest change that addresses it, and check that nothing else broke.
Only if [EXPECTED] is given:
- Describe each failure precisely: what was expected, what happened, and the exact part of the output that is wrong. If no expectation is given and it is not obvious, infer it and say so.
- Group failures into patterns.
- For each pattern, test these causes against the evidence and name the most likely root cause:
- The instruction is missing, ambiguous, or only implied.
- Instructions conflict, or one buried late or deep is outweighed by an earlier one.
- Examples are being copied (length, wording, labels) or do not cover the failing case.
- Input is not delimited, so the model treats data as instructions or mixes it into the answer.
- The output format is underspecified, or the reasoning and the final answer are mixed.
- Missing context or knowledge, so the model fills gaps by guessing.
- Too many jobs in one prompt.
- Not a prompt problem: a capability limit (exact counting, long arithmetic, very long inputs), missing retrieval or tools, settings such as temperature or maximum length, or the pipeline around the model.
- Propose the smallest targeted fix for each root cause, show it as a before and after, and say which failures it should fix and what it might break.
- Give the revised prompt with all fixes applied and nothing else changed.
- Build a quick test set: every failing input, three to five inputs that worked before (to catch regressions), and two new edge cases, each with a pass condition that can be checked.
- Base every diagnosis on evidence in the outputs or the prompt. If the evidence is too thin to tell causes apart, say so and propose a small experiment that would (for example, remove the examples and rerun).
- Prefer explaining the reason behind a rule over adding emphasis.
- Do not claim a fix works; you cannot run it. Say what result would confirm it.
- If a cause is outside the prompt, say so plainly and recommend the right fix (a tool, retrieval, validation code, a setting) rather than more instructions.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Failure patterns
Table: Failure | Expected | Got | Pattern.
Root causes
One short paragraph per pattern with the evidence.
Fixes
Numbered. Each: Before, After, Fixes which failures, Risk.
Revised prompt
Fenced code block.
Test set
Table: Input | Why it is in the set | Pass condition.
If this does not fix it
The next hypothesis to test, and how.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Prompting and assistants
- category
- Prompt engineering
- level
- Intermediate
- made for
- ML / AI engineer, Software engineer, Product manager, Anyone, personal use
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install diagnose-prompt-failures --target claude-codenpx skills add hermes-hq/hodios-dist --skill diagnose-prompt-failures -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-prompting@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Prompt engineeringPrompt engineer
Prompt engineer who writes clear, testable instructions, iterates against real examples and evals, and avoids model-specific tricks. Use for designing, debugging and maintaining prompts.
prompt-engineerImprove a prompt
Diagnoses why a prompt gives weak or inconsistent results and rewrites it with clear context, task, constraints and output format while keeping its intent. Use on any prompt for any AI assistant.
improve-promptCreate few-shot examples
Builds a small set of diverse, representative few-shot examples for a task, including tricky and negative cases, balanced so the model learns the rule rather than copying surface patterns.
create-few-shot-examplesCompress a prompt
Shortens a long prompt while preserving its behaviour, maps every original instruction to where it now lives, reports the real size reduction and lists test inputs to check nothing changed.
compress-promptAdapt a prompt for a reasoning model
Rewrites a prompt for reasoning-capable models by removing step-by-step micromanagement, stating goals, constraints and success criteria, and keeping the output format exact.
adapt-prompt-for-reasoning-modelBuild a test set for a prompt
Builds a hand-run test set for a prompt with happy, edge and negative inputs, expected behaviour and checkable pass criteria per case, and a scoring sheet to compare prompt versions side by side.
build-prompt-test-set