hermes

Measure UX with SUS and task metrics

Plans a UX benchmark with SUS and task metrics, or scores supplied responses, compares them with norms and reports confidence intervals. Use to track UX across releases.

context

The System Usability Scale is the most widely used standard usability questionnaire, and the most often mis-scored. Teams average the raw 1-to-5 answers, forget that even-numbered items are negatively worded, read 68 as "68 per cent", compare two releases from eight people each without any error margin, and drop task metrics that would explain why the score moved. A credible benchmark uses the same tasks and the same kind of participants each time, scores correctly, and reports uncertainty honestly.

task

Only if [RESPONSES] is given: Score and report this UX benchmark.

responses

Product and scope:

product

If no responses were supplied, write a benchmark plan: the 5 to 8 core tasks with success criteria, metrics (task success, time on task for successful attempts, errors, the Single Ease Question after each task, SUS at the end), sample size per user group (20 or more for a stable benchmark, more to detect small differences between releases), unmoderated versus moderated, how to keep later rounds comparable (same tasks, recruitment criteria, environment and order of questionnaires), and a results template. Then stop.

If responses were supplied:

  1. Check the data. Count respondents. Flag rows with missing items, values outside 1 to 5, and straight-lining (the same answer on all 10 items, which is inconsistent because half the items are negatively worded). Say how each is handled: exclude, or keep and flag. If the item order or wording is not the standard SUS, say the scores may not be comparable with norms.
  2. Score SUS correctly. For each respondent: odd items (1, 3, 5, 7, 9) contribute answer minus 1; even items (2, 4, 6, 8, 10) contribute 5 minus answer; the sum is multiplied by 2.5, giving 0 to 100. Show a per-respondent table so the arithmetic can be checked. Then report the mean, standard deviation, median and range.
  3. Confidence interval. Report the 95 per cent interval for the mean SUS: mean plus or minus t (with n minus 1 degrees of freedom) times SD divided by the square root of n. Show the values used.
  4. Compare with norms. State that across large published datasets the average SUS is about 68, and that this is a score, not a percentage. Place the result relative to that average using the interval (clearly above, around, or below), and mention the Sauro-Lewis curved grading scale as a reference without over-reading the letter grade.
  5. Task metrics (if supplied). Per task: success rate with an adjusted-Wald 95 per cent interval (suitable for small samples); time on task for successful attempts as the geometric mean with an interval computed on log times; mean errors per attempt; mean SEQ if collected. Flag tasks with low success or high time as the likely drivers of the SUS score.
  6. Compare with the earlier benchmark (if given). Report the difference with a 95 per cent interval for the difference (Welch's t for independent samples, or a paired comparison if the same people took part). If the interval includes zero, say there is no clear evidence of change. Check that the rounds are comparable before comparing.
  7. With more than about 40 respondents, compute the summary statistics, show the first 10 rows of the per-respondent table, and give a spreadsheet formula for the rest, for example with items in columns B to K: =((B2-1)+(5-C2)+(D2-1)+(5-E2)+(F2-1)+(5-G2)+(H2-1)+(5-I2)+(J2-1)+(5-K2))*2.5.
constraints
  • Never average raw item answers as a score, and never present SUS as a percentage or a percentile.
  • Show your working for every computed figure, and recompute any total you are unsure of. Do not round until the final figures (one decimal place).
  • Do not invent norms, competitor scores or earlier results. With fewer than about 12 respondents, say the interval is wide and the score is indicative only.
  • SUS measures perceived usability overall; it does not say what to fix. Use task data and observations for that.
  • For formal hypothesis testing beyond these intervals, or high-stakes decisions, recommend review by a statistician.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

For a plan: ## Method, ## Tasks (table: task, success criterion, metric), ## Sample and recruitment, ## Results template. For scored data:

Summary

Three to five sentences: the SUS mean with its interval, where it sits against the average, the weakest tasks, and whether it changed since the last round.

Method

SUS results

Per-respondent table (respondent, item contributions, SUS), then summary statistics and the interval.

Task results

| Task | n | Success (95% CI) | Geo-mean time, s (95% CI) | Errors | SEQ |

Comparison

Data quality

Next steps

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Design
category
UX research
level
Intermediate
made for
UX researcher, Product / UX / UI designer, Product manager, Data analyst
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install measure-ux-with-sus --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill measure-ux-with-sus -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the design plugin
claude plugin install hodios-design@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of UX research
PromptUX research

Write a usability test plan

Writes a usability test plan with scenario tasks, success metrics, participant criteria, a screener and a moderator script, tied to the research questions. Use before running a usability study.

write-usability-test-plan
PromptUX research

Synthesize usability test findings

Turns raw usability session notes into evidence-backed issues rated by severity and frequency, with task results and recommendations. Use after a round of usability sessions.

synthesize-usability-findings
PersonaUX research

UX researcher

UX researcher who matches the method to the question, separates what people did from what it means, and protects participants. Use as a partner for planning, running and synthesising research.

ux-researcher
PromptUX research

Analyse session recordings and heatmaps

Synthesises notes from session recordings and heatmaps into usability issues with frequency, severity and evidence, keeping observation apart from interpretation, and plans follow-ups.

analyze-session-recordings
PromptUX research

Build a user journey map from research

Builds an evidence-based journey map with stages, actions, thoughts, emotions, pain points and opportunities, marking every assumption. Use after interviews or studies about one segment.

build-user-journey-map
PromptUX research

Build evidence-based user personas

Builds UX personas from research notes, grouping participants by behaviour, with goals, pain points and scenarios traced to evidence and assumptions marked. Use after interviews or field research.

build-user-personas