Reduce LLM costs and latency
Cuts an LLM feature's cost and latency through prompt trimming, caching, model routing, batching and output limits, each paired with the quality check that proves nothing regressed.
LLM bills usually grow from a few causes: input tokens repeated on every call (long system prompts, tool definitions, full chat history, too many retrieved chunks), a large model used for every request including easy ones, output longer than anyone reads, retries and duplicate calls, and real-time calls for work that could wait. Most savings are safe, but some quietly lower quality, which no one notices until users do. Every change therefore needs a check that would catch a regression before it ships.
Reduce the cost and latency of this feature:
Usage data: Only if [QUALITY_BAR] is given:
Quality bar:
- Build the cost model from the data: calls per user action, input tokens split by part (system prompt, tool definitions, history, retrieved context, user input), output tokens, cached tokens, retries, and the model behind each call. Show which parts make up most of the spend and most of the latency. Where a split is not in the data, estimate it from the sample request and label it as an estimate.
- Generate candidate changes from these levers, keeping only the ones the data supports:
- Remove waste: duplicate or unnecessary calls, retries on non-retryable errors, unused tool definitions, dead instructions.
- Prompt caching: reorder prompts so the stable part (instructions, tool definitions, fixed documents) comes first and the variable part last, then enable the provider's prompt caching. Check the provider's minimum cacheable length and cache lifetime against the traffic pattern.
- Trim context: fewer or better retrieved chunks, history summarised or windowed, shorter instructions that say the same thing.
- Limit output: a maximum output length, a compact format (structured output instead of prose when a program reads it), no restating the input.
- Route by difficulty: send easy requests to a smaller, faster model and escalate on low confidence or failed validation; say how a request is classified.
- Batch: move work that does not need an immediate answer to the provider's batch interface or an off-peak queue.
- Cache responses: exact-match caching for repeated requests; semantic caching only where a near-duplicate answer is acceptable.
- Fine-tuning or distillation into a smaller model: last, only if the eval shows the smaller model cannot reach the bar with prompting.
- For each change, estimate the saving with the arithmetic shown (tokens times calls times price), its effect on latency, the quality risk (none, low, medium, high), and the effort.
- Pair each change with the quality check that must pass before it ships: an offline run on the eval set with a threshold derived from the quality bar, a side-by-side comparison on sampled real traffic, or a shadow or A/B rollout with the metric to watch. If no eval set exists, make building a small one the first change and explain why.
- Order the changes by saving per unit of quality risk and effort, and give a rollout sequence that changes one thing at a time so each saving and each regression can be attributed.
If prices are not in the usage data, do not quote any: use symbols (price per million input tokens, and so on) and show the formula. If the usage data is too thin to find where the money goes, say what to measure first and how.
- Never recommend a change that lowers quality without naming the risk and the check. "Use a cheaper model" alone is not a recommendation.
- Do not invent numbers. Every saving traces back to the usage data or an estimate labelled as one.
- Name providers only as examples; describe caching, batching and routing in general terms with what to check in the provider's documentation.
- Keep user-facing behaviour the same unless the change is listed as a product decision for the owner.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Where the money goes
Table: component | tokens per call | calls per day | share of cost | share of latency.
Ranked changes
Table: # | change | estimated monthly saving | latency effect | quality risk | effort.
Change details
One subsection per change: what to do, the arithmetic, and the quality check with its pass threshold.
Rollout
Numbered order, one change at a time, with the metric to watch after each.
Monitoring
The cost, latency and quality metrics to track per request and the alert thresholds.
Missing data
What would sharpen the estimates and how to collect it.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Backend engineer, Engineering manager, Software engineer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install reduce-llm-costs --target claude-codenpx skills add hermes-hq/hodios-dist --skill reduce-llm-costs -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringWrite an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
write-llm-eval-suiteDesign an LLM agent architecture
Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.
design-agent-architectureReduce cloud spend
Analyses a cloud bill or cost export alongside the architecture and ranks savings by monthly impact, effort and risk. Use when the cloud bill grows faster than usage.
reduce-cloud-spendBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extractionChoose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
choose-ml-approach