hermes

Write an eval suite for an LLM feature

Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.

context

An eval suite is the executable spec of an LLM feature. Without one, every prompt or model change is judged by a few hand-picked examples and regressions ship silently. Suites go wrong in predictable ways: cases that only cover the happy path, a single average score that hides a failing slice, a model judge with a vague rubric that rewards long or confident answers, and thresholds nobody agreed on. Model judges also show position bias and self-preference, so they must be anchored with a rubric and checked against human labels before anyone trusts them.

task

Write an eval suite for this feature: Only if [SAMPLE_INPUTS] is given:

Sample inputs:

Grading approach: .

  1. Turn the feature into success criteria: observable properties of one output that a grader can decide. Mark each as a hard requirement (must hold on every case, such as valid JSON, no leaked system prompt, refusal of out-of-scope requests) or a quality criterion (scored). If the description does not say what a good output is, ask before writing cases.
  2. Write 20 to 40 cases, each tagged with a slice:
  • golden (about 60%): typical inputs, built from the samples when given;
  • edge: empty or minimal input, very long input, mixed languages, ambiguous requests, unusual formatting, boundary values;
  • adversarial: prompt injection inside the user content, requests to reveal instructions, out-of-scope or disallowed requests that fit this feature, inputs designed to trigger the known failure modes. Use invented data only. Give a reference output or the key facts the output must contain wherever one exists.
  1. Pick a grader for each criterion. Use exact match, regex or schema validation for deterministic properties. Use a rubric for qualities. For a model judge, write the judge prompt: the criterion, a 1-to-5 or pass/fail scale with an anchor example for each level, the reference answer when there is one, reasoning before the verdict, and, for pairwise comparisons, both orderings. Say how to calibrate the judge: 20 to 50 human-labelled cases and the agreement level required before it is trusted.
  2. If the grading approach is exact or rubric only, say which criteria it cannot grade reliably and what you would use instead.
  3. Set thresholds: hard requirements at 100%, a pass rate per quality criterion, a minimum per slice, the number of runs per case to absorb sampling variance, and the rule for comparing a candidate against the current version.
constraints
  • Every case must test something a criterion names. Drop cases that duplicate another case's purpose.
  • Do not use real names, emails or customer data in cases.
  • Keep the judge prompt self-contained, so it runs without this conversation.
  • Thresholds are starting values. Say how to revise them after the first runs.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Success criteria

Table: id | criterion | hard or quality | grader.

Cases

One fenced YAML block. Each case: id, slice, input, reference (or must_include), criteria (ids).

Graders

The deterministic checks, the rubric, and the full judge prompt in a fenced block, plus the calibration procedure.

Thresholds and gating

Pass rules per criterion and slice, runs per case, and when a change may ship.

Gaps

What the suite does not cover yet and what data would close it.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
AI and ML engineering
level
Intermediate
made for
ML / AI engineer, Software engineer, QA / test engineer, Product manager
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install write-llm-eval-suite --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill write-llm-eval-suite -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptAI and ML engineering

Design a RAG pipeline

Designs a retrieval-augmented generation pipeline from a corpus and its real questions, covering chunking, hybrid retrieval, reranking, citations and evals. Use before building or rebuilding RAG.

design-rag-pipeline
PromptAI and ML engineering

Plan a fine-tuning project

Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.

plan-fine-tuning
PersonaAI and ML engineering

Machine-learning engineer

Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.

ml-engineer
PromptAI and ML engineering

Build an MCP server

Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.

build-mcp-server
PromptAI and ML engineering

Build an LLM structured extraction step

Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.

build-structured-extraction
PromptAI and ML engineering

Choose between rules, ML and an LLM

Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.

choose-ml-approach