hermes

Build a test set for a prompt

Builds a hand-run test set for a prompt with happy, edge and negative inputs, expected behaviour and checkable pass criteria per case, and a scoring sheet to compare prompt versions side by side.

context

Most prompt changes are judged by running one or two inputs and eyeballing the result, so a fix for one case silently breaks three others. A small fixed test set changes that: every version runs on the same inputs and is scored against the same written criteria. A useful set covers the common case (most of real traffic), edge cases (empty, very long, ambiguous, mixed-language or oddly formatted input, boundary values), and negative cases (input the prompt should refuse, redirect, or answer with "not enough information"). Each case needs an expected behaviour written before running, and a pass criterion someone else could check the same way: an exact match or pattern where possible, a short rubric where judgement is needed.

task

Build a test set of cases for this prompt.

prompt

Only if [REAL_INPUTS] is given:

real inputs

  1. If the prompt's purpose or expected output cannot be worked out, ask one question and stop.
  2. List what the prompt must do: each requirement in it (format, length, content rules, refusal or ask rules, tone), numbered as R1, R2 and so on, plus implicit requirements a user would expect, labelled as implicit.
  3. Plan coverage: about half happy-path cases spread across the realistic variety of inputs, about a third edge cases, and the rest negative cases. Make sure every requirement is exercised by at least one case.
  4. Write each case with a full, realistic input (not a description of an input) for every placeholder. Base cases on the real inputs where given, varied rather than copied; mark synthetic ones.
  5. For each case, write the expected behaviour and a pass criterion, choosing the cheapest reliable check: exact value, contains or does-not-contain, regex, length limit, valid JSON or schema, or a one-sentence rubric for a judge or human.
constraints
  • Inputs must be complete and runnable as written. No "[insert long text here]"; if a long input is needed, write a realistic one or describe exactly how to build it, and flag it.
  • Use fictional names, companies and data; no real personal data.
  • Pass criteria must be specific to this prompt's requirements. Not "the output is good" or "the output is helpful".
  • Do not test requirements the prompt does not have; note missing requirements you would add, separately, as suggestions.
  • If is too small to cover every requirement, say which requirements are untested.
  • The set is meant to be run by hand and scored in the sheet. If the prompt powers a product feature that needs automated graders, thresholds and CI gating, say so in one line and note that these cases can seed that suite.
output format

What it must do

Numbered requirements (R1…), with implicit ones labelled.

Coverage

A small table: Type | Count | Requirements covered.

Test cases

For each case: a heading with ID and short name, then Type, Requirements, Input (in a fenced block, one per placeholder), Expected behaviour, Pass criterion, Check type.

Scoring sheet

A table with one row per case: ID | v1 pass? | v2 pass? | Notes, ready to copy into a spreadsheet.

How to compare versions

Four or five bullets: same settings, several runs per case for variable outputs, compare pass counts per type, read every newly failing case, and do not adopt a version that breaks a negative case.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Prompting and assistants
category
Prompt engineering
level
Intermediate
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install build-prompt-test-set --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill build-prompt-test-set -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the prompting plugin
claude plugin install hodios-prompting@hodios

The plugin brings every entry in this domain at once.

PromptPrompt engineering

Write an LLM-as-judge prompt

Writes an LLM-as-judge grading prompt with a calibrated scale, anchored examples for each score, ordered criteria and a structured verdict, plus checks for common judge biases.

write-judge-prompt
PromptPrompt engineering

Improve a prompt

Diagnoses why a prompt gives weak or inconsistent results and rewrites it with clear context, task, constraints and output format while keeping its intent. Use on any prompt for any AI assistant.

improve-prompt
PromptPrompt engineering

Diagnose prompt failures

Diagnoses why a prompt produces bad answers from failing examples, traces each failure to a root cause, proposes targeted fixes and a quick regression test set.

diagnose-prompt-failures
PromptPrompt engineering

Red-team a prompt

Tests a prompt or assistant setup against adversarial inputs - injection, edge cases, off-topic and harmful requests, data leaks - predicts failures and proposes fixes. For assistant builders.

red-team-prompt
PromptAI and ML engineering

Write an eval suite for an LLM feature

Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.

write-llm-eval-suite
WorkflowPrompt engineering

Prompt iteration track

Improves a prompt in gated steps - define success, build test cases, run and grade, diagnose failures, revise, then compare versions on the same cases before adopting the change.

prompt-iteration-track