Write an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
An eval suite is the executable spec of an LLM feature. Without one, every prompt or model change is judged by a few hand-picked examples and regressions ship silently. Suites go wrong in predictable ways: cases that only cover the happy path, a single average score that hides a failing slice, a model judge with a vague rubric that rewards long or confident answers, and thresholds nobody agreed on. Model judges also show position bias and self-preference, so they must be anchored with a rubric and checked against human labels before anyone trusts them.
Write an eval suite for this feature: Only if [SAMPLE_INPUTS] is given:
Sample inputs:
Grading approach: .
- Turn the feature into success criteria: observable properties of one output that a grader can decide. Mark each as a hard requirement (must hold on every case, such as valid JSON, no leaked system prompt, refusal of out-of-scope requests) or a quality criterion (scored). If the description does not say what a good output is, ask before writing cases.
- Write 20 to 40 cases, each tagged with a slice:
- golden (about 60%): typical inputs, built from the samples when given;
- edge: empty or minimal input, very long input, mixed languages, ambiguous requests, unusual formatting, boundary values;
- adversarial: prompt injection inside the user content, requests to reveal instructions, out-of-scope or disallowed requests that fit this feature, inputs designed to trigger the known failure modes. Use invented data only. Give a reference output or the key facts the output must contain wherever one exists.
- Pick a grader for each criterion. Use exact match, regex or schema validation for deterministic properties. Use a rubric for qualities. For a model judge, write the judge prompt: the criterion, a 1-to-5 or pass/fail scale with an anchor example for each level, the reference answer when there is one, reasoning before the verdict, and, for pairwise comparisons, both orderings. Say how to calibrate the judge: 20 to 50 human-labelled cases and the agreement level required before it is trusted.
- If the grading approach is exact or rubric only, say which criteria it cannot grade reliably and what you would use instead.
- Set thresholds: hard requirements at 100%, a pass rate per quality criterion, a minimum per slice, the number of runs per case to absorb sampling variance, and the rule for comparing a candidate against the current version.
- Every case must test something a criterion names. Drop cases that duplicate another case's purpose.
- Do not use real names, emails or customer data in cases.
- Keep the judge prompt self-contained, so it runs without this conversation.
- Thresholds are starting values. Say how to revise them after the first runs.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Success criteria
Table: id | criterion | hard or quality | grader.
Cases
One fenced YAML block. Each case: id, slice, input, reference (or must_include), criteria (ids).
Graders
The deterministic checks, the rubric, and the full judge prompt in a fenced block, plus the calibration procedure.
Thresholds and gating
Pass rules per criterion and slice, runs per case, and when a change may ship.
Gaps
What the suite does not cover yet and what data would close it.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Software engineer, QA / test engineer, Product manager
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install write-llm-eval-suite --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-llm-eval-suite -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringDesign a RAG pipeline
Designs a retrieval-augmented generation pipeline from a corpus and its real questions, covering chunking, hybrid retrieval, reranking, citations and evals. Use before building or rebuilding RAG.
design-rag-pipelinePlan a fine-tuning project
Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.
plan-fine-tuningMachine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
ml-engineerBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extractionChoose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
choose-ml-approach