hermes

Plan a fine-tuning project

Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.

context

Fine-tuning changes how a model behaves: output format and style, consistency on a narrow classification or extraction task, reliability at calling tools, or a large model's skill distilled into a smaller, cheaper one. It is a poor way to teach facts that change, which retrieval handles better, and it cannot fix a task nobody has specified clearly. Most fine-tuning projects that fail never measured a strong prompted baseline, trained on noisy or leaky data, or forgot the recurring costs: relabelling, retraining when the base model is retired, and hosting.

task

Task:

Data available: Only if [BUDGET] is given:

Budget:

  1. Compare the options for this task: a better prompt with few-shot examples and structured output, retrieval, supervised fine-tuning, preference tuning (only if pairwise preferences exist or can be collected), and distillation from a larger model. Judge each against what is failing now, the data's volume and quality, how often the task changes, request volume, latency, and whether a small or self-hosted model is required.
  2. Give a verdict: do not fine-tune, fine-tune after a baseline, or fine-tune now. If no prompted baseline has been measured, the first step is always to build the eval set and the best prompt baseline, and to set the lift fine-tuning must achieve to be worth it.
  3. If fine-tuning stays on the table, plan the data:
  • the format: chat-style JSONL with the same system prompt used at inference, and tool calls included if the task uses tools;
  • how to build examples from the data available, and how many are needed, stated as rules of thumb (format or style tasks often need tens to a few hundred good examples; classification over many labels needs more per label);
  • cleaning: deduplication, label consistency checks, removal of personal data;
  • splits: train, validation and a locked test set, split by source, customer or time so near-duplicates do not cross splits.
  1. Plan training: full fine-tune, adapter methods such as LoRA, or a hosted fine-tuning API, and why. Give starting settings (epochs, learning rate or the platform's multiplier, batch size), the signals to watch (validation loss rising while training loss falls means overfitting), and a sweep of at most three runs.
  2. Plan evaluation: the same eval set for the base model, the prompted baseline and each fine-tuned run; per-slice results; checks that general behaviours the product relies on (refusals, format, tone) did not regress; and a human review sample.
  3. Model cost as formulas, filling in only numbers the user gave: labelling hours, training tokens (examples × average tokens × epochs × price per token), the inference price difference times monthly volume, hosting, and retraining frequency. Give the break-even volume.
  4. State go/no-go criteria and how to roll back.
constraints
  • Never invent prices or benchmark results. Use variables where the user gave no figure.
  • Keep the plan vendor-neutral. Name a platform only as an example.
  • If the budget cannot cover the plan, say what to cut first.
  • Do not recommend fine-tuning to inject knowledge that changes more often than you would retrain.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Verdict

One line, then two or three sentences of reasoning.

Why

Table: approach | fit for this task | cost | main risk.

Baseline first

The prompt baseline to build, the eval set, and the target lift.

Data plan

Format, sources, cleaning, splits and target size.

Training plan

Method, starting settings, runs and what to watch.

Evaluation

What is compared, on which slices, and what counts as a win.

Cost model

One-off and recurring costs as formulas, with break-even volume.

Go/no-go

The criteria to ship, and the rollback.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
AI and ML engineering
level
Expert
made for
ML / AI engineer, Data scientist, Software engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install plan-fine-tuning --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill plan-fine-tuning -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptAI and ML engineering

Choose between rules, ML and an LLM

Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.

choose-ml-approach
PromptAI and ML engineering

Write an eval suite for an LLM feature

Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.

write-llm-eval-suite
PromptAI and ML engineering

Review a training dataset sample

Audits a sample of a labelled dataset for label noise, leakage, duplicates, class imbalance and representation gaps, and gives a concrete fix for each problem. Use before training or fine-tuning.

review-training-data
PersonaAI and ML engineering

Machine-learning engineer

Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.

ml-engineer
PromptAI and ML engineering

Build an MCP server

Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.

build-mcp-server
PromptAI and ML engineering

Build an LLM structured extraction step

Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.

build-structured-extraction