Plan a caching strategy
Designs caching for a slow path, covering what to cache at which layer, keys, TTLs, invalidation, stampede protection and measuring hit rate and staleness. Use when fixing latency or database load.
Caching is the fastest way to make a slow path fast and one of the easiest ways to make a system wrong. Common failures: caching before finding why the path is slow (a missing index would have fixed it), keys that leak one user's data to another because the user or tenant was not in the key, invalidation that misses a write path so stale data lives forever, every entry expiring at once and stampeding the database, a cache outage taking the whole service down because nothing could serve without it, and no metric that shows whether the cache helps. A good plan caches only where it pays, states the staleness each layer allows, and is measured.
Design caching for this slow path:
Freshness needs:
Only if [STACK] is given: Stack: Only if [TRAFFIC] is given: Traffic:
- Decide first whether caching is the right fix. If the evidence points to an unindexed query, an N+1 pattern, a chatty remote call or an algorithmic problem, say so and recommend fixing that first or alongside. If there is no measurement of where time goes, say what to measure before building anything.
- Choose the layers, from closest to the user outwards, and say what each caches and why: HTTP caching with Cache-Control and ETags, CDN or edge caching (only for content that is public or correctly varied), application-level shared cache (for example Redis or Memcached), in-process memory cache (small, hot, rarely changing data, and only with an invalidation story for multiple instances), database-level options (materialised views, read replicas), and memoisation of expensive computations. Use only the layers that pay.
- Define keys: include every input that changes the result (tenant, user or permission scope, locale, currency, query parameters, feature flags, schema or code version), normalise inputs to avoid duplicate entries, and put a version prefix in the key so a deploy can invalidate safely. Call out any layer where personal or permission-dependent data could be served to the wrong user.
- Define TTLs and invalidation per data type, mapped to the freshness needs: cache-aside with TTL, write-through, explicit invalidation or event-driven invalidation on writes, or stale-while-revalidate. For explicit invalidation, list every write path that must trigger it and the race between a write and a concurrent cache fill (and how to avoid it, for example deleting after commit, or versioned values). Add TTL jitter so entries do not expire together.
- Protect against stampedes and failures: request coalescing or a per-key lock for refills, early probabilistic refresh or serving stale while one request refreshes, negative caching for "not found" with a short TTL, a size limit and eviction policy, timeouts on cache calls, and graceful degradation when the cache is down (fall back to the source with load shedding, never fail the request just because the cache failed).
- Estimate the benefit with arithmetic from the traffic numbers: expected hit rate given the access skew, the load removed from the source, memory needed (entries × average size), and latency at the expected hit rate. Mark assumed numbers.
- Define measurement: hit and miss rate per key family, latency for hits and misses, source load before and after, evictions, memory use, and a staleness check (for example sampling cached values against the source).
- Give a rollout plan: behind a flag, one key family at a time, with the success criteria and how to turn it off.
- If the stack is known, include a short code sketch of the cache-aside read with stampede protection for the main key family.
- Never cache responses that depend on the user's identity or permissions in a shared layer without the identity or scope in the key, and never in a public CDN.
- Every cached item must have a TTL, even when it is also invalidated explicitly.
- Respect the stated freshness needs exactly. If a need cannot be met with caching, say so.
- Do not invent current latency, hit rates or traffic numbers; mark assumptions.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Verdict
Whether caching is the right fix, what else to fix first, and the expected benefit, in at most 5 lines.
Cache plan
Table: layer, what is cached, why, staleness allowed.
Keys and TTLs
Table: key family, key format, TTL with jitter, size estimate.
Invalidation
Per key family: strategy, the write paths that trigger it, and the race handling.
Failure and stampede handling
Bullets.
Measurement
Table: metric, target, alert.
Rollout
Numbered steps, then the code sketch if the stack is known.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Performance
- level
- Intermediate
- made for
- Backend engineer, Software engineer, Software architect, Site reliability engineer
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install plan-caching-strategy --target claude-codenpx skills add hermes-hq/hodios-dist --skill plan-caching-strategy -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of PerformancePerformance engineer
Acts as a performance engineer who profiles before optimising, changes one thing at a time and reports gains with numbers and variance. Use for latency, throughput or memory work.
performance-engineerBackend engineer
Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.
backend-engineerProfile and speed up a hot path
Measures a slow operation, profiles where the time goes, and makes it faster one verified change at a time, with before-and-after numbers. Use when an endpoint, command or function is too slow.
profile-hot-pathOptimise a slow SQL query
Speeds up a slow SQL query from its execution plan, proposing rewrites and indexes with expected gains and their write-cost trade-offs. Use when one query dominates latency or database load.
optimize-sql-queryFix N+1 queries
Finds N+1 database queries behind an endpoint, page or job by counting real queries, fixes them with eager loading or batching, and adds a query-count test so they do not return.
fix-n-plus-one-queriesPlan a load test
Designs a load test with a workload model, scenarios, ramp profile and pass or fail thresholds, then writes the script for the chosen tool. Use before a launch, a traffic event or a capacity decision.
plan-load-test