hermes

Reduce LLM costs and latency

Cuts an LLM feature's cost and latency through prompt trimming, caching, model routing, batching and output limits, each paired with the quality check that proves nothing regressed.

context

LLM bills usually grow from a few causes: input tokens repeated on every call (long system prompts, tool definitions, full chat history, too many retrieved chunks), a large model used for every request including easy ones, output longer than anyone reads, retries and duplicate calls, and real-time calls for work that could wait. Most savings are safe, but some quietly lower quality, which no one notices until users do. Every change therefore needs a check that would catch a regression before it ships.

task

Reduce the cost and latency of this feature:

Usage data: Only if [QUALITY_BAR] is given:

Quality bar:

  1. Build the cost model from the data: calls per user action, input tokens split by part (system prompt, tool definitions, history, retrieved context, user input), output tokens, cached tokens, retries, and the model behind each call. Show which parts make up most of the spend and most of the latency. Where a split is not in the data, estimate it from the sample request and label it as an estimate.
  2. Generate candidate changes from these levers, keeping only the ones the data supports:
  • Remove waste: duplicate or unnecessary calls, retries on non-retryable errors, unused tool definitions, dead instructions.
  • Prompt caching: reorder prompts so the stable part (instructions, tool definitions, fixed documents) comes first and the variable part last, then enable the provider's prompt caching. Check the provider's minimum cacheable length and cache lifetime against the traffic pattern.
  • Trim context: fewer or better retrieved chunks, history summarised or windowed, shorter instructions that say the same thing.
  • Limit output: a maximum output length, a compact format (structured output instead of prose when a program reads it), no restating the input.
  • Route by difficulty: send easy requests to a smaller, faster model and escalate on low confidence or failed validation; say how a request is classified.
  • Batch: move work that does not need an immediate answer to the provider's batch interface or an off-peak queue.
  • Cache responses: exact-match caching for repeated requests; semantic caching only where a near-duplicate answer is acceptable.
  • Fine-tuning or distillation into a smaller model: last, only if the eval shows the smaller model cannot reach the bar with prompting.
  1. For each change, estimate the saving with the arithmetic shown (tokens times calls times price), its effect on latency, the quality risk (none, low, medium, high), and the effort.
  2. Pair each change with the quality check that must pass before it ships: an offline run on the eval set with a threshold derived from the quality bar, a side-by-side comparison on sampled real traffic, or a shadow or A/B rollout with the metric to watch. If no eval set exists, make building a small one the first change and explain why.
  3. Order the changes by saving per unit of quality risk and effort, and give a rollout sequence that changes one thing at a time so each saving and each regression can be attributed.

If prices are not in the usage data, do not quote any: use symbols (price per million input tokens, and so on) and show the formula. If the usage data is too thin to find where the money goes, say what to measure first and how.

constraints
  • Never recommend a change that lowers quality without naming the risk and the check. "Use a cheaper model" alone is not a recommendation.
  • Do not invent numbers. Every saving traces back to the usage data or an estimate labelled as one.
  • Name providers only as examples; describe caching, batching and routing in general terms with what to check in the provider's documentation.
  • Keep user-facing behaviour the same unless the change is listed as a product decision for the owner.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Where the money goes

Table: component | tokens per call | calls per day | share of cost | share of latency.

Ranked changes

Table: # | change | estimated monthly saving | latency effect | quality risk | effort.

Change details

One subsection per change: what to do, the arithmetic, and the quality check with its pass threshold.

Rollout

Numbered order, one change at a time, with the metric to watch after each.

Monitoring

The cost, latency and quality metrics to track per request and the alert thresholds.

Missing data

What would sharpen the estimates and how to collect it.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
AI and ML engineering
level
Intermediate
made for
ML / AI engineer, Backend engineer, Engineering manager, Software engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install reduce-llm-costs --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill reduce-llm-costs -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptAI and ML engineering

Write an eval suite for an LLM feature

Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.

write-llm-eval-suite
PromptAI and ML engineering

Design an LLM agent architecture

Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.

design-agent-architecture
PromptDevOps

Reduce cloud spend

Analyses a cloud bill or cost export alongside the architecture and ranks savings by monthly impact, effort and risk. Use when the cloud bill grows faster than usage.

reduce-cloud-spend
PromptAI and ML engineering

Build an MCP server

Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.

build-mcp-server
PromptAI and ML engineering

Build an LLM structured extraction step

Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.

build-structured-extraction
PromptAI and ML engineering

Choose between rules, ML and an LLM

Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.

choose-ml-approach