hermes

Choose between rules, ML and an LLM

Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.

context

Two defaults waste the most money. Sending every request to a large LLM is slow and costly at volume, and hard to test when the logic is really a dozen rules. Training a custom model when there are fifty examples and the requirements change monthly wastes weeks. The right choice depends on a few facts: whether the logic can be written down, how variable the input is, how much labelled data exists, the cost of an error, volume and latency, explainability requirements, how often the task changes, and who will maintain the result. Hybrids are often best: rules for the clear cases with a model for the rest, or an LLM to label data that then trains a small, cheap model.

task

Recommend an approach for: Only if [DATA_AVAILABLE] is given:

Data available: Only if [CONSTRAINTS] is given:

Constraints:

  1. Restate the problem as input, output, volume, latency budget and cost of an error. If volume, latency or labelled data is missing and could flip the recommendation, ask for it. Otherwise state an assumption and continue.
  2. Evaluate each option against this problem, not in general:
  • rules or heuristics (including regular expressions, lookups and templates);
  • classical ML (logistic regression, gradient-boosted trees, small text classifiers) on engineered features;
  • a hosted LLM with prompting, few-shot examples and structured output;
  • a fine-tuned or distilled model;
  • the hybrids that fit.
  1. For each option, reason about the accuracy you can expect and why, cost per thousand requests as a formula or order of magnitude with stated assumptions, latency, the data required, maintenance work, failure modes and explainability.
  2. Recommend one approach, give the cheapest experiment that would confirm it within days, and name the observations that should make the team switch.
constraints
  • Show the reasoning that connects each fact about the problem to the recommendation.
  • Do not invent accuracy figures. Give expectations as ranges to verify, and say what they rest on.
  • Never recommend fine-tuning before a prompted baseline has been measured, or an LLM where a lookup table would do.
  • Prefer the option the team can run and debug, all else being equal.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Recommendation

One paragraph: the approach and the two or three facts that decide it.

Problem as stated

Input, output, volume, latency, cost of an error, with assumptions marked.

Comparison

Table: option | expected accuracy | cost per 1,000 | latency | data needed | maintenance | main failure mode.

Validation experiment

The smallest test that would confirm the choice, and its pass bar.

Switch triggers

What would make you change approach, and to what.

Assumptions

Every number or fact you supplied yourself.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
AI and ML engineering
level
Intermediate
made for
ML / AI engineer, Software engineer, Tech lead / staff engineer, Product manager
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install choose-ml-approach --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill choose-ml-approach -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptAI and ML engineering

Plan a fine-tuning project

Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.

plan-fine-tuning
PromptAI and ML engineering

Plan a machine-learning experiment

Plans a machine-learning experiment before any training code exists: framing, baselines, leak-proof splits, metrics, ablations and a stop rule. Use when starting a new model or modelling spike.

plan-ml-experiment
PersonaAI and ML engineering

Machine-learning engineer

Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.

ml-engineer
PromptAI and ML engineering

Build an MCP server

Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.

build-mcp-server
PromptAI and ML engineering

Build an LLM structured extraction step

Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.

build-structured-extraction
PromptAI and ML engineering

Design an LLM agent architecture

Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.

design-agent-architecture