Plan a fine-tuning project
Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.
Fine-tuning changes how a model behaves: output format and style, consistency on a narrow classification or extraction task, reliability at calling tools, or a large model's skill distilled into a smaller, cheaper one. It is a poor way to teach facts that change, which retrieval handles better, and it cannot fix a task nobody has specified clearly. Most fine-tuning projects that fail never measured a strong prompted baseline, trained on noisy or leaky data, or forgot the recurring costs: relabelling, retraining when the base model is retired, and hosting.
Task:
Data available: Only if [BUDGET] is given:
Budget:
- Compare the options for this task: a better prompt with few-shot examples and structured output, retrieval, supervised fine-tuning, preference tuning (only if pairwise preferences exist or can be collected), and distillation from a larger model. Judge each against what is failing now, the data's volume and quality, how often the task changes, request volume, latency, and whether a small or self-hosted model is required.
- Give a verdict: do not fine-tune, fine-tune after a baseline, or fine-tune now. If no prompted baseline has been measured, the first step is always to build the eval set and the best prompt baseline, and to set the lift fine-tuning must achieve to be worth it.
- If fine-tuning stays on the table, plan the data:
- the format: chat-style JSONL with the same system prompt used at inference, and tool calls included if the task uses tools;
- how to build examples from the data available, and how many are needed, stated as rules of thumb (format or style tasks often need tens to a few hundred good examples; classification over many labels needs more per label);
- cleaning: deduplication, label consistency checks, removal of personal data;
- splits: train, validation and a locked test set, split by source, customer or time so near-duplicates do not cross splits.
- Plan training: full fine-tune, adapter methods such as LoRA, or a hosted fine-tuning API, and why. Give starting settings (epochs, learning rate or the platform's multiplier, batch size), the signals to watch (validation loss rising while training loss falls means overfitting), and a sweep of at most three runs.
- Plan evaluation: the same eval set for the base model, the prompted baseline and each fine-tuned run; per-slice results; checks that general behaviours the product relies on (refusals, format, tone) did not regress; and a human review sample.
- Model cost as formulas, filling in only numbers the user gave: labelling hours, training tokens (examples × average tokens × epochs × price per token), the inference price difference times monthly volume, hosting, and retraining frequency. Give the break-even volume.
- State go/no-go criteria and how to roll back.
- Never invent prices or benchmark results. Use variables where the user gave no figure.
- Keep the plan vendor-neutral. Name a platform only as an example.
- If the budget cannot cover the plan, say what to cut first.
- Do not recommend fine-tuning to inject knowledge that changes more often than you would retrain.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Verdict
One line, then two or three sentences of reasoning.
Why
Table: approach | fit for this task | cost | main risk.
Baseline first
The prompt baseline to build, the eval set, and the target lift.
Data plan
Format, sources, cleaning, splits and target size.
Training plan
Method, starting settings, runs and what to watch.
Evaluation
What is compared, on which slices, and what counts as a win.
Cost model
One-off and recurring costs as formulas, with break-even volume.
Go/no-go
The criteria to ship, and the rollback.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Expert
- made for
- ML / AI engineer, Data scientist, Software engineer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install plan-fine-tuning --target claude-codenpx skills add hermes-hq/hodios-dist --skill plan-fine-tuning -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringChoose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
choose-ml-approachWrite an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
write-llm-eval-suiteReview a training dataset sample
Audits a sample of a labelled dataset for label noise, leakage, duplicates, class imbalance and representation gaps, and gives a concrete fix for each problem. Use before training or fine-tuning.
review-training-dataMachine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
ml-engineerBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extraction