hermes

Plan validation of a survey scale or test

Plans the validation of a survey scale or test, covering content and construct validity, reliability, factor analysis, invariance, sample sizes and reporting. For researchers building measures.

context

Validity is not a property of an instrument but of the interpretation and use of its scores in a given population (Standards for Educational and Psychological Testing; COSMIN for health measures). A validation plan therefore starts from the intended use and builds an argument from several kinds of evidence: content (do items cover the construct, judged by experts and the target group), response process (cognitive interviews), internal structure (factor analysis, dimensionality, measurement invariance), relations to other variables (convergent, discriminant, known-groups, criterion), and consequences. Reliability (internal consistency, test-retest, inter-rater) is necessary but not sufficient. Common mistakes: treating a high Cronbach's alpha as proof of validity, running exploratory and confirmatory factor analysis on the same sample, using fit-index cut-offs as strict rules, skipping invariance before comparing groups, and translating a scale without cultural adaptation.

task

Plan the validation of this instrument.

instrument

Only if [POPULATION] is given: Population:

  1. State the intended use and the score interpretations to be supported (for example "rank individuals", "compare groups", "detect change over time", "screen against a cut-off"). Each claim determines the evidence needed.
  2. Design the validation in phases with the evidence each provides: construct definition and item review; content validity with an expert panel (for example item and scale content validity indices) and target-group review; cognitive interviews; pilot and item analysis; structural validation; reliability; relations to other variables; invariance across the subgroups that will be compared; and responsiveness or cut-off derivation if the use requires them. Skip phases that do not apply and say why.
  3. For an adapted or translated instrument, add forward and back translation, reconciliation, harmonisation and cognitive testing, and plan invariance testing against the original language version where data allow.
  4. Give a sample size plan per phase with reasoning, not a single rule of thumb: for factor analysis, account for the number of items and factors, expected communalities and the need for separate samples (or a split sample) for exploratory and confirmatory analysis; for test-retest, the precision of the intraclass correlation and a retest interval matched to how stable the construct is; for known-groups and convergent analyses, the expected effect or correlation.
  5. Specify the analyses: item distributions, floor and ceiling effects, item-total correlations; estimator choice for ordinal items (for example WLSMV or polychoric-based estimation); fit indices reported together with their conventional benchmarks and the caveat that they are guides, not pass marks; reliability with McDonald's omega alongside alpha and their confidence intervals; ICC model and type for test-retest; standard error of measurement and smallest detectable change if change will be measured; configural, metric and scalar invariance; and where item response theory would add value.
  6. List what to report and which guideline or checklist to follow.
constraints
  • Name convergent and discriminant measures only as types ("an established measure of loneliness"), unless the user named them; never invent a validated instrument or its psychometric properties.
  • Say plainly when the intended use is not supportable by the plan (for example clinical screening without a reference standard).
  • Treat numeric benchmarks as conventions with a source tradition, not laws, and say how to proceed if the data miss them (re-specify with theory, not with modification indices alone).
  • If the construct definition or items are missing or unclear, list exactly what is needed and give the plan with stated assumptions rather than guessing the items.
output format

Intended use and claims

The use, the score interpretations, and the evidence each needs.

Validation plan

A table: phase | purpose | method | participants | output | evidence type.

Sample size plan

Per phase, with reasoning.

Analysis plan

Bulleted by phase, with the decision rules.

Reporting checklist

The items to report and the guideline that applies.

Risks and open questions

What could undermine validity and what you need from the user.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Research and science
category
Research methods
level
Expert
made for
Researcher / scientist, Student, UX researcher, Data scientist
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-03
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install validate-measurement-instrument --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill validate-measurement-instrument -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the research-science plugin
claude plugin install hodios-research-science@hodios

The plugin brings every entry in this domain at once.

PromptResearch methods

Write a survey questionnaire

Writes an unbiased questionnaire for a stated research aim, with construct mapping, appropriate response scales, skip logic and a pilot checklist. Use before fielding any survey.

write-survey-questionnaire
PromptResearch methods

Design a research study

Designs a study from a research question to hypotheses, design, variables, sampling and sample size, a pre-specified analysis plan and threats to validity. Use before collecting any data.

design-research-study
PromptResearch methods

Write a study preregistration

Writes a study preregistration with hypotheses, design, sampling plan, variables, exclusion rules and a pre-specified analysis plan, flagging open choices. For OSF or AsPredicted users.

write-preregistration
PersonaResearch methods

Research methodologist

Research methodologist who probes study designs for validity threats, matches methods to questions and asks what evidence would change the conclusion. Use as a sparring partner for any study.

research-methodologist
PersonaStatistics

Consulting statistician

Consulting statistician who asks how the data were produced before analysing them, chooses methods that fit the question, checks assumptions and refuses to over-claim. Use for any data analysis.

statistician
PromptResearch methods

Build a qualitative codebook

Builds a codebook and coding procedure for qualitative data, with definitions, inclusion and exclusion rules, verbatim examples and an inter-rater check. Use before coding interviews or open text.

build-qualitative-codebook