hermes

Write an LLM-as-judge prompt

Writes an LLM-as-judge grading prompt with a calibrated scale, anchored examples for each score, ordered criteria and a structured verdict, plus checks for common judge biases.

context

A model grading another model's output is useful only if its scores agree with careful human judgement. Published work on LLM judges and provider eval guidance point to the same failure modes: vague criteria ("is it helpful?"), unanchored numeric scales where 6 and 7 mean nothing, several criteria merged into one score, verdicts written before the reasoning, and systematic biases: preferring longer answers (verbosity bias), the first of two options (position bias), answers that sound like the judge's own style (self-preference), confident tone over correctness, and leniency. Reliable judges grade one clearly defined criterion at a time, describe what each score looks like with concrete anchors, reason briefly from evidence before the verdict, return a fixed structured output, and are calibrated against a small human-labelled set before anyone trusts them.

task

Write a judge prompt on a scale.

task and good output

Only if [CRITERIA] is given:

criteria

  1. If there is no way to tell what a good output is (no task description or no good example), ask for one and stop.
  2. Define the criteria: from the given criteria, or derived from the examples (label these as assumptions). Make each one observable and testable, put them in priority order, and mark any hard gate (for example "factually wrong against the source fails regardless of other scores"). Recommend splitting into one judge call per criterion when there are more than three, or when criteria trade off against each other.
  3. Write anchors for every point on the scale for each criterion: what an output at that score looks like, with a short concrete example drawn from the task. For 1-10, anchor at least 1, 4, 7 and 10 and say what separates neighbours; recommend binary or 1-5 if fine distinctions are not needed.
  4. Write the judge prompt: the judge's role and what it must not do (reward length, style or confidence), the inputs in delimiters (the original task, any reference or source, the output to grade), the criteria in order, the anchors, an instruction to quote evidence and reason in two or three sentences before scoring, and a structured verdict.
  5. List bias checks and how to run them, and a calibration plan.
constraints
  • The verdict must be machine-readable: JSON with, per criterion, evidence (short quote), reasoning (at most three sentences), and score, then an overall field defined by an explicit rule (for example "fail if any gate fails, otherwise the mean").
  • Instruct the judge to grade only against the criteria and reference given, to treat "I don't know" or a refusal according to an explicit rule, and to score an output the same regardless of length beyond what the criteria require.
  • For pairwise comparison, require running both orders and counting only consistent preferences.
  • Do not invent ground truth: if correctness needs a reference answer or source, add a slot for it in the judge prompt.
  • Keep the judge prompt model-agnostic and under about 700 words.
output format

Criteria

A table: # | Criterion | Definition | Gate? (yes/no). Then any assumptions.

Judge prompt

The full judge prompt in one fenced block, including the anchors and the JSON verdict schema.

Bias checks

A table: Bias | How to test it | Mitigation in this prompt.

Calibration plan

Numbered steps: label 30 to 50 outputs by hand, run the judge, measure agreement (percent agreement for binary, a rank or kappa statistic for scales), read every disagreement, adjust anchors, and re-run; the agreement level to reach before relying on it.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Prompting and assistants
category
Prompt engineering
level
Expert
made for
ML / AI engineer, Product manager, Researcher / scientist
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install write-judge-prompt --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill write-judge-prompt -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the prompting plugin
claude plugin install hodios-prompting@hodios

The plugin brings every entry in this domain at once.

PromptPrompt engineering

Build a test set for a prompt

Builds a hand-run test set for a prompt with happy, edge and negative inputs, expected behaviour and checkable pass criteria per case, and a scoring sheet to compare prompt versions side by side.

build-prompt-test-set
PromptAI and ML engineering

Write an eval suite for an LLM feature

Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.

write-llm-eval-suite
PromptPrompt engineering

Diagnose prompt failures

Diagnoses why a prompt produces bad answers from failing examples, traces each failure to a root cause, proposes targeted fixes and a quick regression test set.

diagnose-prompt-failures
PersonaPrompt engineering

Prompt engineer

Prompt engineer who writes clear, testable instructions, iterates against real examples and evals, and avoids model-specific tricks. Use for designing, debugging and maintaining prompts.

prompt-engineer
PromptPrompt engineering

Improve a prompt

Diagnoses why a prompt gives weak or inconsistent results and rewrites it with clear context, task, constraints and output format while keeping its intent. Use on any prompt for any AI assistant.

improve-prompt
PromptPrompt engineering

Adapt a prompt for a reasoning model

Rewrites a prompt for reasoning-capable models by removing step-by-step micromanagement, stating goals, constraints and success criteria, and keeping the output format exact.

adapt-prompt-for-reasoning-model