Build a test set for a prompt
Builds a hand-run test set for a prompt with happy, edge and negative inputs, expected behaviour and checkable pass criteria per case, and a scoring sheet to compare prompt versions side by side.
Most prompt changes are judged by running one or two inputs and eyeballing the result, so a fix for one case silently breaks three others. A small fixed test set changes that: every version runs on the same inputs and is scored against the same written criteria. A useful set covers the common case (most of real traffic), edge cases (empty, very long, ambiguous, mixed-language or oddly formatted input, boundary values), and negative cases (input the prompt should refuse, redirect, or answer with "not enough information"). Each case needs an expected behaviour written before running, and a pass criterion someone else could check the same way: an exact match or pattern where possible, a short rubric where judgement is needed.
Build a test set of cases for this prompt.
Only if [REAL_INPUTS] is given:
- If the prompt's purpose or expected output cannot be worked out, ask one question and stop.
- List what the prompt must do: each requirement in it (format, length, content rules, refusal or ask rules, tone), numbered as R1, R2 and so on, plus implicit requirements a user would expect, labelled as implicit.
- Plan coverage: about half happy-path cases spread across the realistic variety of inputs, about a third edge cases, and the rest negative cases. Make sure every requirement is exercised by at least one case.
- Write each case with a full, realistic input (not a description of an input) for every placeholder. Base cases on the real inputs where given, varied rather than copied; mark synthetic ones.
- For each case, write the expected behaviour and a pass criterion, choosing the cheapest reliable check: exact value, contains or does-not-contain, regex, length limit, valid JSON or schema, or a one-sentence rubric for a judge or human.
- Inputs must be complete and runnable as written. No "[insert long text here]"; if a long input is needed, write a realistic one or describe exactly how to build it, and flag it.
- Use fictional names, companies and data; no real personal data.
- Pass criteria must be specific to this prompt's requirements. Not "the output is good" or "the output is helpful".
- Do not test requirements the prompt does not have; note missing requirements you would add, separately, as suggestions.
- If is too small to cover every requirement, say which requirements are untested.
- The set is meant to be run by hand and scored in the sheet. If the prompt powers a product feature that needs automated graders, thresholds and CI gating, say so in one line and note that these cases can seed that suite.
What it must do
Numbered requirements (R1…), with implicit ones labelled.
Coverage
A small table: Type | Count | Requirements covered.
Test cases
For each case: a heading with ID and short name, then Type, Requirements, Input (in a fenced block, one per placeholder), Expected behaviour, Pass criterion, Check type.
Scoring sheet
A table with one row per case: ID | v1 pass? | v2 pass? | Notes, ready to copy into a spreadsheet.
How to compare versions
Four or five bullets: same settings, several runs per case for variable outputs, compare pass counts per type, read every newly failing case, and do not adopt a version that breaks a negative case.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Prompting and assistants
- category
- Prompt engineering
- level
- Intermediate
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-03
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install build-prompt-test-set --target claude-codenpx skills add hermes-hq/hodios-dist --skill build-prompt-test-set -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-prompting@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Prompt engineeringWrite an LLM-as-judge prompt
Writes an LLM-as-judge grading prompt with a calibrated scale, anchored examples for each score, ordered criteria and a structured verdict, plus checks for common judge biases.
write-judge-promptImprove a prompt
Diagnoses why a prompt gives weak or inconsistent results and rewrites it with clear context, task, constraints and output format while keeping its intent. Use on any prompt for any AI assistant.
improve-promptDiagnose prompt failures
Diagnoses why a prompt produces bad answers from failing examples, traces each failure to a root cause, proposes targeted fixes and a quick regression test set.
diagnose-prompt-failuresRed-team a prompt
Tests a prompt or assistant setup against adversarial inputs - injection, edge cases, off-topic and harmful requests, data leaks - predicts failures and proposes fixes. For assistant builders.
red-team-promptWrite an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
write-llm-eval-suitePrompt iteration track
Improves a prompt in gated steps - define success, build test cases, run and grade, diagnose failures, revise, then compare versions on the same cases before adopting the change.
prompt-iteration-track