hermes

Plan a machine-learning experiment

Plans a machine-learning experiment before any training code exists: framing, baselines, leak-proof splits, metrics, ablations and a stop rule. Use when starting a new model or modelling spike.

context

Weeks of modelling are lost to the same mistakes. With no baseline, "0.92 AUC" means nothing. Random splits on data with time or group structure leak the answer into training. Features computed after the moment of prediction make offline results impossible to reproduce in production. The chosen metric does not match the decision the model supports. Tuning against the test set inflates every number. Without a stop rule, the project drifts from run to run. All of this is cheapest to fix on paper, before training code exists.

task

Plan an experiment for this problem:

Dataset: Only if [COMPUTE_BUDGET] is given:

Compute budget:

  1. Frame it: the target, the unit of prediction (a row, user, session, document), the moment of prediction and which features exist at that moment, the decision the output drives, and the cost of a false positive against a false negative. If the target or the moment of prediction is unclear, ask before planning further.
  2. Choose metrics: one primary metric that matches the decision (for example recall at a fixed precision for rare positives, PR-AUC for imbalanced ranking, MAE in the target's units), guardrail metrics, the slices to report separately, and the smallest improvement that would change the decision.
  3. Define baselines in order: a trivial one (majority class, mean, last value, seasonal naive), a heuristic a domain expert would write, and a simple model such as logistic regression or gradient-boosted trees on obvious features. Every later result is reported against all three.
  4. Design the splits: by time when the model will predict the future, by group when the same user, patient or document appears in many rows, stratified when classes are rare, cross-validated when data is small. Lock the test set until the final evaluation.
  5. List leakage checks specific to this dataset: features recorded after the moment of prediction, identifiers or timestamps that correlate with the label, duplicates or near-duplicates across splits, preprocessing fitted on all the data, and target encoding computed outside the training fold. For each, give the concrete check, and treat a result that looks too good as a leak until proven otherwise.
  6. Write the run plan: ordered runs, each with a hypothesis, the single change, its expected effect, its compute cost, and the evidence that would confirm it. Include ablations that attribute any gain over the simple model, and at least three seeds wherever variance could exceed the gain.
  7. Specify reproducibility: data snapshot or version, code commit, configuration and seeds recorded for every run.
  8. Write the stop rule: the condition to stop (target met, budget spent, or no gain above the minimum over a set number of consecutive runs) and the result that would end the project.
constraints
  • Do not write training code. This is the plan the code will follow.
  • Fit the run plan inside the compute budget, and say what to drop if it does not fit.
  • Prefer the simplest model that meets the decision's needs. A complex model must beat the simple one by more than seed variance to stay in the plan.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Framing

Target, unit, moment of prediction, decision, error costs.

Metrics

Primary, guardrails, slices, minimum meaningful improvement.

Baselines

The three baselines and how each is computed.

Data splits

The split scheme and why it matches how the model will be used.

Leakage checks

Checklist: suspected leak, check, action if found.

Run plan

Table: # | hypothesis | change | cost | what confirms it.

Reproducibility

What is recorded for every run, and where.

Stop rule

When to stop, and what would end the project.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
AI and ML engineering
level
Intermediate
made for
ML / AI engineer, Data scientist, Researcher / scientist
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install plan-ml-experiment --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill plan-ml-experiment -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptAI and ML engineering

Review a training dataset sample

Audits a sample of a labelled dataset for label noise, leakage, duplicates, class imbalance and representation gaps, and gives a concrete fix for each problem. Use before training or fine-tuning.

review-training-data
PromptAI and ML engineering

Choose between rules, ML and an LLM

Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.

choose-ml-approach
PersonaAI and ML engineering

Machine-learning engineer

Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.

ml-engineer
PromptAI and ML engineering

Build an MCP server

Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.

build-mcp-server
PromptAI and ML engineering

Build an LLM structured extraction step

Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.

build-structured-extraction
PromptAI and ML engineering

Design an LLM agent architecture

Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.

design-agent-architecture