Plan a load test
Designs a load test with a workload model, scenarios, ramp profile and pass or fail thresholds, then writes the script for the chosen tool. Use before a launch, a traffic event or a capacity decision.
Most load tests answer the wrong question. They hammer one endpoint with a fixed number of looping users, hit only cached data, and report an average latency. Closed-model loops slow down when the system slows down, which hides the very saturation the test was meant to find (coordinated omission). A useful test models real arrival rates and the real mix of requests, uses varied data, and ends with a clear pass or fail against agreed thresholds.
Design a load test for: Only if [TRAFFIC_PROFILE] is given: Production traffic:
- State the objective as a question the test answers, such as "Does checkout meet p95 below 800 ms at 2x last Black Friday peak?". If the system description does not reveal the question, ask and stop.
- Build the workload model: an open model with arrival rates (requests or iterations per second) for user-facing traffic; the mix of transactions by weight; think time; test data variety large enough to defeat caches the way real traffic does; authentication handling. If traffic data is missing, propose numbers, label them assumptions, and say how to derive the real ones from access logs.
- Define scenarios: a smoke test, load at expected peak, a stress test that ramps past peak to find the breaking point, a spike, and a soak of several hours when leaks or slow degradation are a concern. Give the ramp for each.
- Set pass and fail thresholds: latency percentiles (p95 and p99, never only the average), error rate, and the throughput achieved versus the target. List the server-side saturation signals to watch (CPU, memory, connection pools, queue depth, database load).
- Write the script for implementing the model, the scenarios and the thresholds as automatic pass or fail where the tool supports it.
- Never point the test at production or at third-party services (payment providers, email, SMS) without explicit approval; stub or sandbox them and say so.
- Check the load generator itself is not the bottleneck, and say how.
- Exclude warm-up from the results.
- Do not invent endpoints or payloads; use placeholders where the description has none and list them.
Objective
The question, and the decision it informs.
Workload model
A table: transaction, share of traffic, target rate at peak, think time, test data source.
Scenarios
A table: scenario, ramp, duration, purpose.
Pass and fail criteria
A table: metric, threshold, source (client or server).
Script
One fenced block for , followed by any placeholders to fill.
Run checklist
Environment parity, data reset, monitoring in place, people to notify, and how to abort.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Performance
- level
- Intermediate
- made for
- Backend engineer, Site reliability engineer, QA / test engineer
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install plan-load-test --target claude-codenpx skills add hermes-hq/hodios-dist --skill plan-load-test -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of PerformanceDefine SLOs and burn-rate alerts
Defines SLIs, SLOs and an error-budget policy from a service's user journeys, with multi-window burn-rate alert rules. Use when alerting is noisy or reliability targets are vague.
define-slosFix N+1 queries
Finds N+1 database queries behind an endpoint, page or job by counting real queries, fixes them with eager loading or batching, and adds a query-count test so they do not return.
fix-n-plus-one-queriesProfile and speed up a hot path
Measures a slow operation, profiles where the time goes, and makes it faster one verified change at a time, with before-and-after numbers. Use when an endpoint, command or function is too slow.
profile-hot-pathReduce JavaScript bundle size
Measures a web app's JavaScript bundles, finds the largest avoidable contributors, and shrinks them with verified changes ranked by bytes saved. Use when page load is slow or a size budget is blown.
reduce-bundle-sizeFind a memory leak
Finds a memory leak from heap snapshots, memory metrics and code, naming the retaining path and the minimal fix with a regression check. Use when memory grows until a process is killed or restarted.
find-memory-leakImprove Core Web Vitals
Diagnoses poor Core Web Vitals (LCP, INP, CLS) from a Lighthouse, field-data or trace report and ranks fixes by expected improvement. Use when a page fails the vitals thresholds.
improve-web-vitals