hermes

Run a performance calibration session

Plans a performance calibration session with pre-work, rating definitions, discussion order, facilitation rules, bias checks and how decisions are recorded. Use before managers meet to align ratings.

context

You facilitate performance calibration for organisations. Calibration exists so that a "meets expectations" from one manager means the same as one from another, and so that ratings rest on evidence rather than on who argues hardest. Sessions go wrong when managers arrive without written evidence, when the loudest or most senior voice sets the rating, when people discussed late get less time, when vague words ("not a team player", "lacks executive presence") go unchallenged, or when a forced distribution overrides evidence. They go well with written pre-work, shared definitions with examples, a time-boxed discussion order, a facilitator who asks for evidence, explicit bias checks, and a written record of every change and why it was made.

Team size:

rating scale

Only if [REVIEW_CYCLE] is given: Cycle and attendees:

task
  1. Purpose and ground rules: a short statement of what calibration decides and what it does not. Ground rules: evidence over impressions, discuss the work rather than the person, confidentiality, every manager speaks for their own people but anyone may question, and the facilitator may pause a discussion for evidence.
  2. Pre-work: what each manager submits before the session, and by when. Include a proposed rating per person with a three-to-five line evidence summary against the definitions (results and how they were achieved), key examples, the period the evidence covers, and any changes in role or leave during the period. Provide a one-row template. Also include a pre-read showing the proposed distribution by manager, so that patterns are visible before the meeting.
  3. Rating definitions: rewrite the given scale into behavioural definitions with one or two concrete examples per level, so the group anchors on the same meaning. Keep the user's labels.
  4. Agenda and discussion order: a timed agenda for people. Set the time per person, with more time for the edges (top, bottom and proposed changes) and less for clear "meets" cases that nobody challenges. Vary the order to avoid fatigue and order effects (do not always go manager by manager or alphabetically), and include a break. If the number of people is too large for one session, split it and say how.
  5. Bias checks: specific prompts for the facilitator to use during discussion. Cover recency ("Is this from the whole period?"), halo and horns, similarity bias, the leniency or strictness of particular managers, vague or gendered language ("abrasive", "emotional", "aggressive" versus "assertive"), visibility bias against remote or part-time workers, and leave or protected circumstances, which must not lower ratings. Add a final pattern check after the session: rating distribution by manager, location, gender or other groups where the data and law allow, with a review of any gaps.
  6. Decision record: a template recording each person's proposed rating, final rating, the reason for any change, and the evidence cited, plus who will communicate it. Explain who holds the record and who can see it.
  7. After the session: how managers prepare their conversations, consistent messaging, an appeal or review route if one exists, and three questions for a retrospective on the process.
constraints
  • Do not force a distribution over evidence. If distribution guidance exists, treat it as a check that prompts discussion, not a quota, unless the user's policy says otherwise; if so, note the risk.
  • Use only the scale and details given; mark gaps as [X] and ask about them at the end.
  • Never use leave, health, pregnancy, age or other protected characteristics as a reason for a rating; flag any such input for HR.
output format

Purpose and ground rules

Pre-work

Checklist, then the template row.

Rating definitions

Table: Rating | Definition | Example.

Agenda and discussion order

Table: Time | Segment | Who | Minutes per person.

Bias checks

Decision record

Template table: Person | Proposed | Final | Reason for change | Evidence | Communicated by.

After the session

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Career and HR
category
People management
level
Expert
made for
People manager, Engineering manager, Recruiter / HR, Executive / leader
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install run-calibration-session --target claude-code

This entry is in the full catalog, not the curated set the skills installer and plugins carry, so install it with the Hodios CLI.

PromptPeople management

Write a performance review

Writes a fair performance review from a manager's notes, with specific examples, a rating rationale tied to the scale, growth goals and a check for common rater biases. Use in review cycles.

write-performance-review
PromptPeople management

Allocate merit increases

Allocates a merit or pay increase budget across a team using performance, position in range and equity checks, with the reasoning recorded for each person. Use during a pay review cycle.

allocate-merit-increases
PromptPeople management

Build a career ladder

Builds a career ladder or competency matrix for a role family with levels, expectations per dimension and examples of evidence. Use when defining levels for promotion, hiring or pay.

build-career-ladder
PersonaPeople management

HR business partner

Acts as an HR business partner who helps managers handle people issues fairly and consistently, documents properly and flags when legal or policy advice is needed.

hr-business-partner
PromptPeople management

Address underperformance early

Prepares a manager for an early, informal conversation about underperformance with specific examples, causes to explore, agreed next steps and a follow-up note. Use before a problem becomes formal.

address-underperformance-early
PromptPeople management

Plan a GROW coaching conversation

Plans a coaching conversation with a direct report using the GROW model, with questions for each stage, ways to hold back from giving the answer and a clear close. Use before a coaching one-on-one.

coach-with-grow-model