hermes

Add rate limiting to an API

Adds rate limiting to API endpoints with a fitting algorithm, keys, per-tier limits, standard headers, 429 responses and tests. Use when protecting endpoints from abuse or overload.

context

Rate limiting goes wrong in a few repeatable ways: limits keyed by client IP when every request arrives from the load balancer's address, or keyed by a spoofable X-Forwarded-For; login limits keyed by account and IP together, which a botnet rotating IPs walks straight past, or a hard per-account lockout that lets anyone lock a victim out; a limiter that blocks every login when its store goes down; in-memory counters on six instances that quietly allow six times the limit; a read-then-write counter in Redis that races under load; fixed windows that allow double the limit at the window boundary; 429 responses with no hint of when to retry, so clients hammer harder; and limits switched on in production without anyone knowing which customers they would block. Good rate limiting picks the key and algorithm per purpose, is atomic, tells clients what is happening and is rolled out in observe-only mode first.

task

Add rate limiting to these endpoints:

endpoints

Only if [TRAFFIC_PROFILE] is given: Traffic profile: Only if [STACK] is given: Stack: Counter storage: (auto: in-memory only for a single instance, otherwise the shared store the app already runs; ask before adding a new one)

  1. Read the app's middleware chain, auth, proxy configuration, existing rate limiting (including at a gateway, CDN or WAF) and how many instances run. Do not add a second limiter on top of an existing one without saying why.
  2. Define the policy per endpoint group, in a table:
  • Purpose: abuse prevention (login, sign-up, password reset, OTP), fair use per customer, or overload protection.
  • Key: authenticated user or API key for fair use. For login, password reset and OTP endpoints, two independent limits: one per target account identifier across all IPs (stops guessing one account from many IPs; slow it with growing delays or a challenge rather than a hard lockout an attacker can trigger on purpose) and one per client IP across all accounts (stops one source spraying many accounts). Client IP only when there is no identity, always derived from the trusted proxy hop (configure the framework's trusted-proxy setting rather than reading the header blindly). Say plainly that per-IP limits do not stop distributed credential stuffing, and name what complements them (breached-password checks, bot management at the CDN, MFA).
  • Algorithm: token bucket or GCRA when bursts are acceptable, sliding window (log or counter) when the limit must be smooth; avoid plain fixed windows unless the boundary burst is acceptable, and say so.
  • Limits: per tier or plan, with burst size. Propose numbers from the traffic profile with the reasoning, marked as proposed if no profile was given.
  1. Implement it with the framework's middleware or a well-maintained library already in use or common for the stack. With a shared store, make the check-and-increment atomic (a single atomic command or a server-side script), set expiry on every key, and decide what happens when the store is unavailable: fail open for fair-use limits; for login-style endpoints fall back to a stricter per-instance in-memory limit rather than rejecting every login, which would turn a cache outage into an auth outage. Log and emit a metric either way.
  2. Respond correctly: HTTP 429 with a Retry-After header, a consistent error body in the API's existing error format, and rate-limit headers on responses. Use the RateLimit-Policy and RateLimit header fields from the IETF HTTPAPI draft if the API has no existing convention, or the widely used X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset if clients already expect those; say which and why.
  3. Add allowlisting for health checks and internal callers where needed, and make limits configurable without a deploy.
  4. Add observability: a metric of allowed and limited requests by endpoint group and tier, and a log line for limited requests with the key hashed or truncated.
  5. Write tests with a fake or controllable clock: requests under the limit pass, the limit plus one returns 429 with Retry-After, the bucket refills over time, different keys do not interfere, tiers get their own limits, the spoofed X-Forwarded-For case does not bypass the limit, and the store-down behaviour matches the chosen policy. Run them and report the real result.
  6. Recommend a rollout: log-only (shadow) mode first, review who would have been limited, then enforce.
constraints
  • Do not use in-memory counters when there is more than one instance unless the limit is explicitly per instance; say so if it is.
  • Never key on a client-supplied header without a trusted-proxy configuration.
  • Keep limits and tier names in configuration, not hard-coded in handlers.
  • Do not claim a header draft is a final standard; describe it as the IETF draft.
  • Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
  • Keep the change as small as it can be while still being correct.
  • Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
  • If you could not run a check, say so plainly and say which one.
output format

Policy

Table: endpoint group, purpose, key or keys, algorithm, limit and burst per tier, store-down behaviour. Proposed numbers are marked proposed.

Design

Where the limiter sits in the request path, the storage and atomicity approach, and the headers, in a few bullets.

Changes

One line per file.

Tests

One line per test and the real result of the run.

Rollout

Numbered steps from shadow mode to enforcement, with what to watch.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Implementation
level
Intermediate
made for
Backend engineer, Software engineer, Site reliability engineer
needs
repo-read, file-write, shell
risk
runs-commands
version
v1.1.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install add-rate-limiting --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill add-rate-limiting -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Implementation
PersonaImplementation

Backend engineer

Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.

backend-engineer
PromptImplementation

Build a REST endpoint end to end

Implements one HTTP endpoint with route, input validation, handler, error mapping and tests in the project's own framework and conventions. Use when adding an API route.

build-rest-endpoint
PromptSecurity

Harden web app headers and cookies

Produces hardened HTTP security headers, a Content Security Policy, CORS and cookie settings for a web app, rolled out first in report-only mode. Use before launch or after a security scan.

harden-web-app-config
PromptPerformance

Plan a load test

Designs a load test with a workload model, scenarios, ramp profile and pass or fail thresholds, then writes the script for the chosen tool. Use before a launch, a traffic event or a capacity decision.

plan-load-test
PromptImplementation

Put a change behind a feature flag

Wraps new behaviour behind a feature flag with a safe default, a kill switch, tests for both paths and a cleanup ticket. Use when shipping a risky change incrementally.

add-feature-flag
PromptImplementation

Build a reusable UI component

Builds a typed, accessible UI component from a description or screenshot, with loading, empty and error states and a usage example. Use when adding a component to a frontend.

build-ui-component