hermes

Plan a caching strategy

Designs caching for a slow path, covering what to cache at which layer, keys, TTLs, invalidation, stampede protection and measuring hit rate and staleness. Use when fixing latency or database load.

context

Caching is the fastest way to make a slow path fast and one of the easiest ways to make a system wrong. Common failures: caching before finding why the path is slow (a missing index would have fixed it), keys that leak one user's data to another because the user or tenant was not in the key, invalidation that misses a write path so stale data lives forever, every entry expiring at once and stampeding the database, a cache outage taking the whole service down because nothing could serve without it, and no metric that shows whether the cache helps. A good plan caches only where it pays, states the staleness each layer allows, and is measured.

task

Design caching for this slow path:

hot path

Freshness needs:

freshness

Only if [STACK] is given: Stack: Only if [TRAFFIC] is given: Traffic:

  1. Decide first whether caching is the right fix. If the evidence points to an unindexed query, an N+1 pattern, a chatty remote call or an algorithmic problem, say so and recommend fixing that first or alongside. If there is no measurement of where time goes, say what to measure before building anything.
  2. Choose the layers, from closest to the user outwards, and say what each caches and why: HTTP caching with Cache-Control and ETags, CDN or edge caching (only for content that is public or correctly varied), application-level shared cache (for example Redis or Memcached), in-process memory cache (small, hot, rarely changing data, and only with an invalidation story for multiple instances), database-level options (materialised views, read replicas), and memoisation of expensive computations. Use only the layers that pay.
  3. Define keys: include every input that changes the result (tenant, user or permission scope, locale, currency, query parameters, feature flags, schema or code version), normalise inputs to avoid duplicate entries, and put a version prefix in the key so a deploy can invalidate safely. Call out any layer where personal or permission-dependent data could be served to the wrong user.
  4. Define TTLs and invalidation per data type, mapped to the freshness needs: cache-aside with TTL, write-through, explicit invalidation or event-driven invalidation on writes, or stale-while-revalidate. For explicit invalidation, list every write path that must trigger it and the race between a write and a concurrent cache fill (and how to avoid it, for example deleting after commit, or versioned values). Add TTL jitter so entries do not expire together.
  5. Protect against stampedes and failures: request coalescing or a per-key lock for refills, early probabilistic refresh or serving stale while one request refreshes, negative caching for "not found" with a short TTL, a size limit and eviction policy, timeouts on cache calls, and graceful degradation when the cache is down (fall back to the source with load shedding, never fail the request just because the cache failed).
  6. Estimate the benefit with arithmetic from the traffic numbers: expected hit rate given the access skew, the load removed from the source, memory needed (entries × average size), and latency at the expected hit rate. Mark assumed numbers.
  7. Define measurement: hit and miss rate per key family, latency for hits and misses, source load before and after, evictions, memory use, and a staleness check (for example sampling cached values against the source).
  8. Give a rollout plan: behind a flag, one key family at a time, with the success criteria and how to turn it off.
  9. If the stack is known, include a short code sketch of the cache-aside read with stampede protection for the main key family.
constraints
  • Never cache responses that depend on the user's identity or permissions in a shared layer without the identity or scope in the key, and never in a public CDN.
  • Every cached item must have a TTL, even when it is also invalidated explicitly.
  • Respect the stated freshness needs exactly. If a need cannot be met with caching, say so.
  • Do not invent current latency, hit rates or traffic numbers; mark assumptions.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Verdict

Whether caching is the right fix, what else to fix first, and the expected benefit, in at most 5 lines.

Cache plan

Table: layer, what is cached, why, staleness allowed.

Keys and TTLs

Table: key family, key format, TTL with jitter, size estimate.

Invalidation

Per key family: strategy, the write paths that trigger it, and the race handling.

Failure and stampede handling

Bullets.

Measurement

Table: metric, target, alert.

Rollout

Numbered steps, then the code sketch if the stack is known.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Performance
level
Intermediate
made for
Backend engineer, Software engineer, Software architect, Site reliability engineer
risk
read-only
version
v1.0.0 · incubating
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install plan-caching-strategy --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill plan-caching-strategy -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Performance
PersonaPerformance

Performance engineer

Acts as a performance engineer who profiles before optimising, changes one thing at a time and reports gains with numbers and variance. Use for latency, throughput or memory work.

performance-engineer
PersonaImplementation

Backend engineer

Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.

backend-engineer
PromptPerformance

Profile and speed up a hot path

Measures a slow operation, profiles where the time goes, and makes it faster one verified change at a time, with before-and-after numbers. Use when an endpoint, command or function is too slow.

profile-hot-path
PromptPerformance

Optimise a slow SQL query

Speeds up a slow SQL query from its execution plan, proposing rewrites and indexes with expected gains and their write-cost trade-offs. Use when one query dominates latency or database load.

optimize-sql-query
PromptPerformance

Fix N+1 queries

Finds N+1 database queries behind an endpoint, page or job by counting real queries, fixes them with eager loading or batching, and adds a query-count test so they do not return.

fix-n-plus-one-queries
PromptPerformance

Plan a load test

Designs a load test with a workload model, scenarios, ramp profile and pass or fail thresholds, then writes the script for the chosen tool. Use before a launch, a traffic event or a capacity decision.

plan-load-test