Write an LLM-as-judge prompt
Writes an LLM-as-judge grading prompt with a calibrated scale, anchored examples for each score, ordered criteria and a structured verdict, plus checks for common judge biases.
A model grading another model's output is useful only if its scores agree with careful human judgement. Published work on LLM judges and provider eval guidance point to the same failure modes: vague criteria ("is it helpful?"), unanchored numeric scales where 6 and 7 mean nothing, several criteria merged into one score, verdicts written before the reasoning, and systematic biases: preferring longer answers (verbosity bias), the first of two options (position bias), answers that sound like the judge's own style (self-preference), confident tone over correctness, and leniency. Reliable judges grade one clearly defined criterion at a time, describe what each score looks like with concrete anchors, reason briefly from evidence before the verdict, return a fixed structured output, and are calibrated against a small human-labelled set before anyone trusts them.
Write a judge prompt on a scale.
Only if [CRITERIA] is given:
- If there is no way to tell what a good output is (no task description or no good example), ask for one and stop.
- Define the criteria: from the given criteria, or derived from the examples (label these as assumptions). Make each one observable and testable, put them in priority order, and mark any hard gate (for example "factually wrong against the source fails regardless of other scores"). Recommend splitting into one judge call per criterion when there are more than three, or when criteria trade off against each other.
- Write anchors for every point on the scale for each criterion: what an output at that score looks like, with a short concrete example drawn from the task. For 1-10, anchor at least 1, 4, 7 and 10 and say what separates neighbours; recommend binary or 1-5 if fine distinctions are not needed.
- Write the judge prompt: the judge's role and what it must not do (reward length, style or confidence), the inputs in delimiters (the original task, any reference or source, the output to grade), the criteria in order, the anchors, an instruction to quote evidence and reason in two or three sentences before scoring, and a structured verdict.
- List bias checks and how to run them, and a calibration plan.
- The verdict must be machine-readable: JSON with, per criterion,
evidence(short quote),reasoning(at most three sentences), andscore, then anoverallfield defined by an explicit rule (for example "fail if any gate fails, otherwise the mean"). - Instruct the judge to grade only against the criteria and reference given, to treat "I don't know" or a refusal according to an explicit rule, and to score an output the same regardless of length beyond what the criteria require.
- For pairwise comparison, require running both orders and counting only consistent preferences.
- Do not invent ground truth: if correctness needs a reference answer or source, add a slot for it in the judge prompt.
- Keep the judge prompt model-agnostic and under about 700 words.
Criteria
A table: # | Criterion | Definition | Gate? (yes/no). Then any assumptions.
Judge prompt
The full judge prompt in one fenced block, including the anchors and the JSON verdict schema.
Bias checks
A table: Bias | How to test it | Mitigation in this prompt.
Calibration plan
Numbered steps: label 30 to 50 outputs by hand, run the judge, measure agreement (percent agreement for binary, a rank or kappa statistic for scales), read every disagreement, adjust anchors, and re-run; the agreement level to reach before relying on it.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Prompting and assistants
- category
- Prompt engineering
- level
- Expert
- made for
- ML / AI engineer, Product manager, Researcher / scientist
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-03
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install write-judge-prompt --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-judge-prompt -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-prompting@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Prompt engineeringBuild a test set for a prompt
Builds a hand-run test set for a prompt with happy, edge and negative inputs, expected behaviour and checkable pass criteria per case, and a scoring sheet to compare prompt versions side by side.
build-prompt-test-setWrite an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
write-llm-eval-suiteDiagnose prompt failures
Diagnoses why a prompt produces bad answers from failing examples, traces each failure to a root cause, proposes targeted fixes and a quick regression test set.
diagnose-prompt-failuresPrompt engineer
Prompt engineer who writes clear, testable instructions, iterates against real examples and evals, and avoids model-specific tricks. Use for designing, debugging and maintaining prompts.
prompt-engineerImprove a prompt
Diagnoses why a prompt gives weak or inconsistent results and rewrites it with clear context, task, constraints and output format while keeping its intent. Use on any prompt for any AI assistant.
improve-promptAdapt a prompt for a reasoning model
Rewrites a prompt for reasoning-capable models by removing step-by-step micromanagement, stating goals, constraints and success criteria, and keeping the output format exact.
adapt-prompt-for-reasoning-model