Measure UX with SUS and task metrics
Plans a UX benchmark with SUS and task metrics, or scores supplied responses, compares them with norms and reports confidence intervals. Use to track UX across releases.
The System Usability Scale is the most widely used standard usability questionnaire, and the most often mis-scored. Teams average the raw 1-to-5 answers, forget that even-numbered items are negatively worded, read 68 as "68 per cent", compare two releases from eight people each without any error margin, and drop task metrics that would explain why the score moved. A credible benchmark uses the same tasks and the same kind of participants each time, scores correctly, and reports uncertainty honestly.
Only if [RESPONSES] is given: Score and report this UX benchmark.
Product and scope:
If no responses were supplied, write a benchmark plan: the 5 to 8 core tasks with success criteria, metrics (task success, time on task for successful attempts, errors, the Single Ease Question after each task, SUS at the end), sample size per user group (20 or more for a stable benchmark, more to detect small differences between releases), unmoderated versus moderated, how to keep later rounds comparable (same tasks, recruitment criteria, environment and order of questionnaires), and a results template. Then stop.
If responses were supplied:
- Check the data. Count respondents. Flag rows with missing items, values outside 1 to 5, and straight-lining (the same answer on all 10 items, which is inconsistent because half the items are negatively worded). Say how each is handled: exclude, or keep and flag. If the item order or wording is not the standard SUS, say the scores may not be comparable with norms.
- Score SUS correctly. For each respondent: odd items (1, 3, 5, 7, 9) contribute answer minus 1; even items (2, 4, 6, 8, 10) contribute 5 minus answer; the sum is multiplied by 2.5, giving 0 to 100. Show a per-respondent table so the arithmetic can be checked. Then report the mean, standard deviation, median and range.
- Confidence interval. Report the 95 per cent interval for the mean SUS: mean plus or minus t (with n minus 1 degrees of freedom) times SD divided by the square root of n. Show the values used.
- Compare with norms. State that across large published datasets the average SUS is about 68, and that this is a score, not a percentage. Place the result relative to that average using the interval (clearly above, around, or below), and mention the Sauro-Lewis curved grading scale as a reference without over-reading the letter grade.
- Task metrics (if supplied). Per task: success rate with an adjusted-Wald 95 per cent interval (suitable for small samples); time on task for successful attempts as the geometric mean with an interval computed on log times; mean errors per attempt; mean SEQ if collected. Flag tasks with low success or high time as the likely drivers of the SUS score.
- Compare with the earlier benchmark (if given). Report the difference with a 95 per cent interval for the difference (Welch's t for independent samples, or a paired comparison if the same people took part). If the interval includes zero, say there is no clear evidence of change. Check that the rounds are comparable before comparing.
- With more than about 40 respondents, compute the summary statistics, show the first 10 rows of the per-respondent table, and give a spreadsheet formula for the rest, for example with items in columns B to K:
=((B2-1)+(5-C2)+(D2-1)+(5-E2)+(F2-1)+(5-G2)+(H2-1)+(5-I2)+(J2-1)+(5-K2))*2.5.
- Never average raw item answers as a score, and never present SUS as a percentage or a percentile.
- Show your working for every computed figure, and recompute any total you are unsure of. Do not round until the final figures (one decimal place).
- Do not invent norms, competitor scores or earlier results. With fewer than about 12 respondents, say the interval is wide and the score is indicative only.
- SUS measures perceived usability overall; it does not say what to fix. Use task data and observations for that.
- For formal hypothesis testing beyond these intervals, or high-stakes decisions, recommend review by a statistician.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
For a plan: ## Method, ## Tasks (table: task, success criterion, metric), ## Sample and recruitment, ## Results template.
For scored data:
Summary
Three to five sentences: the SUS mean with its interval, where it sits against the average, the weakest tasks, and whether it changed since the last round.
Method
SUS results
Per-respondent table (respondent, item contributions, SUS), then summary statistics and the interval.
Task results
| Task | n | Success (95% CI) | Geo-mean time, s (95% CI) | Errors | SEQ |
Comparison
Data quality
Next steps
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Design
- category
- UX research
- level
- Intermediate
- made for
- UX researcher, Product / UX / UI designer, Product manager, Data analyst
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install measure-ux-with-sus --target claude-codenpx skills add hermes-hq/hodios-dist --skill measure-ux-with-sus -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-design@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of UX researchWrite a usability test plan
Writes a usability test plan with scenario tasks, success metrics, participant criteria, a screener and a moderator script, tied to the research questions. Use before running a usability study.
write-usability-test-planSynthesize usability test findings
Turns raw usability session notes into evidence-backed issues rated by severity and frequency, with task results and recommendations. Use after a round of usability sessions.
synthesize-usability-findingsUX researcher
UX researcher who matches the method to the question, separates what people did from what it means, and protects participants. Use as a partner for planning, running and synthesising research.
ux-researcherAnalyse session recordings and heatmaps
Synthesises notes from session recordings and heatmaps into usability issues with frequency, severity and evidence, keeping observation apart from interpretation, and plans follow-ups.
analyze-session-recordingsBuild a user journey map from research
Builds an evidence-based journey map with stages, actions, thoughts, emotions, pain points and opportunities, marking every assumption. Use after interviews or studies about one segment.
build-user-journey-mapBuild evidence-based user personas
Builds UX personas from research notes, grouping participants by behaviour, with goals, pain points and scenarios traced to evidence and assumptions marked. Use after interviews or field research.
build-user-personas