hermes

Plan a load test

Designs a load test with a workload model, scenarios, ramp profile and pass or fail thresholds, then writes the script for the chosen tool. Use before a launch, a traffic event or a capacity decision.

context

Most load tests answer the wrong question. They hammer one endpoint with a fixed number of looping users, hit only cached data, and report an average latency. Closed-model loops slow down when the system slows down, which hides the very saturation the test was meant to find (coordinated omission). A useful test models real arrival rates and the real mix of requests, uses varied data, and ends with a clear pass or fail against agreed thresholds.

task

Design a load test for: Only if [TRAFFIC_PROFILE] is given: Production traffic:

  1. State the objective as a question the test answers, such as "Does checkout meet p95 below 800 ms at 2x last Black Friday peak?". If the system description does not reveal the question, ask and stop.
  2. Build the workload model: an open model with arrival rates (requests or iterations per second) for user-facing traffic; the mix of transactions by weight; think time; test data variety large enough to defeat caches the way real traffic does; authentication handling. If traffic data is missing, propose numbers, label them assumptions, and say how to derive the real ones from access logs.
  3. Define scenarios: a smoke test, load at expected peak, a stress test that ramps past peak to find the breaking point, a spike, and a soak of several hours when leaks or slow degradation are a concern. Give the ramp for each.
  4. Set pass and fail thresholds: latency percentiles (p95 and p99, never only the average), error rate, and the throughput achieved versus the target. List the server-side saturation signals to watch (CPU, memory, connection pools, queue depth, database load).
  5. Write the script for implementing the model, the scenarios and the thresholds as automatic pass or fail where the tool supports it.
constraints
  • Never point the test at production or at third-party services (payment providers, email, SMS) without explicit approval; stub or sandbox them and say so.
  • Check the load generator itself is not the bottleneck, and say how.
  • Exclude warm-up from the results.
  • Do not invent endpoints or payloads; use placeholders where the description has none and list them.
output format

Objective

The question, and the decision it informs.

Workload model

A table: transaction, share of traffic, target rate at peak, think time, test data source.

Scenarios

A table: scenario, ramp, duration, purpose.

Pass and fail criteria

A table: metric, threshold, source (client or server).

Script

One fenced block for , followed by any placeholders to fill.

Run checklist

Environment parity, data reset, monitoring in place, people to notify, and how to abort.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Performance
level
Intermediate
made for
Backend engineer, Site reliability engineer, QA / test engineer
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install plan-load-test --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill plan-load-test -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Performance
PromptIncident and operations

Define SLOs and burn-rate alerts

Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.

define-slos
PromptPerformance

Fix N+1 queries

Finds N+1 database queries behind an endpoint, page or job by counting real queries, fixes them with eager loading or batching, and adds a query-count test so they do not return.

fix-n-plus-one-queries
PromptPerformance

Profile and speed up a hot path

Measures a slow operation, profiles where the time goes, and makes it faster one verified change at a time, with before-and-after numbers. Use when an endpoint, command or function is too slow.

profile-hot-path
PromptPerformance

Reduce JavaScript bundle size

Measures a web app's JavaScript bundles, finds the largest avoidable contributors, and shrinks them with verified changes ranked by bytes saved. Use when page load is slow or a size budget is blown.

reduce-bundle-size
PromptPerformance

Find a memory leak

Finds a memory leak from heap snapshots, memory metrics and code, naming the retaining path and the minimal fix with a regression check. Use when memory grows until a process is killed or restarted.

find-memory-leak
PromptPerformance

Improve Core Web Vitals

Diagnoses poor Core Web Vitals (LCP, INP, CLS) from a Lighthouse, field-data or trace report and ranks fixes by expected improvement. Use when a page fails the vitals thresholds.

improve-web-vitals