Choose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
Two defaults waste the most money. Sending every request to a large LLM is slow and costly at volume, and hard to test when the logic is really a dozen rules. Training a custom model when there are fifty examples and the requirements change monthly wastes weeks. The right choice depends on a few facts: whether the logic can be written down, how variable the input is, how much labelled data exists, the cost of an error, volume and latency, explainability requirements, how often the task changes, and who will maintain the result. Hybrids are often best: rules for the clear cases with a model for the rest, or an LLM to label data that then trains a small, cheap model.
Recommend an approach for: Only if [DATA_AVAILABLE] is given:
Data available: Only if [CONSTRAINTS] is given:
Constraints:
- Restate the problem as input, output, volume, latency budget and cost of an error. If volume, latency or labelled data is missing and could flip the recommendation, ask for it. Otherwise state an assumption and continue.
- Evaluate each option against this problem, not in general:
- rules or heuristics (including regular expressions, lookups and templates);
- classical ML (logistic regression, gradient-boosted trees, small text classifiers) on engineered features;
- a hosted LLM with prompting, few-shot examples and structured output;
- a fine-tuned or distilled model;
- the hybrids that fit.
- For each option, reason about the accuracy you can expect and why, cost per thousand requests as a formula or order of magnitude with stated assumptions, latency, the data required, maintenance work, failure modes and explainability.
- Recommend one approach, give the cheapest experiment that would confirm it within days, and name the observations that should make the team switch.
- Show the reasoning that connects each fact about the problem to the recommendation.
- Do not invent accuracy figures. Give expectations as ranges to verify, and say what they rest on.
- Never recommend fine-tuning before a prompted baseline has been measured, or an LLM where a lookup table would do.
- Prefer the option the team can run and debug, all else being equal.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Recommendation
One paragraph: the approach and the two or three facts that decide it.
Problem as stated
Input, output, volume, latency, cost of an error, with assumptions marked.
Comparison
Table: option | expected accuracy | cost per 1,000 | latency | data needed | maintenance | main failure mode.
Validation experiment
The smallest test that would confirm the choice, and its pass bar.
Switch triggers
What would make you change approach, and to what.
Assumptions
Every number or fact you supplied yourself.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Software engineer, Tech lead / staff engineer, Product manager
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install choose-ml-approach --target claude-codenpx skills add hermes-hq/hodios-dist --skill choose-ml-approach -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringPlan a fine-tuning project
Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.
plan-fine-tuningPlan a machine-learning experiment
Plans a machine-learning experiment before any training code exists: framing, baselines, leak-proof splits, metrics, ablations and a stop rule. Use when starting a new model or modelling spike.
plan-ml-experimentMachine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
ml-engineerBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extractionDesign an LLM agent architecture
Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.
design-agent-architecture