Review a training dataset sample
Audits a sample of a labelled dataset for label noise, leakage, duplicates, class imbalance and representation gaps, and gives a concrete fix for each problem. Use before training or fine-tuning.
A model cannot be more consistent than its labels. Most dataset problems are systematic: a guideline that two labellers read differently, a source whose rows are all one class, a field that leaks the label, or thousands of near-identical rows that inflate test scores. Reading a sample row by row finds these problems far more cheaply than training a model and wondering why it plateaus. The aim is to find the patterns behind individual errors, not to relabel the sample.
Audit this sample for the task below.
Task:
Sample:
- Identify what one row represents, which column is the label, and the label set. If the label column or a label's meaning is unclear, ask before auditing.
- Read every row and check for:
- label noise: rows whose label contradicts their content. Separate clear errors from ambiguous rows that reveal a guideline gap;
- inconsistency: near-identical rows with different labels;
- duplicates and near-duplicates, and across splits if there is a split column;
- leakage: fields or text that give away the label (label words in the text, status tags, identifiers, timestamps recorded after the outcome, boilerplate unique to one source);
- class balance: counts per label in the sample;
- representation gaps: languages, lengths, sources, time periods or user groups that are missing or rare, and the edge cases the task implies but the sample lacks;
- formatting defects: truncation, encoding errors, HTML or template residue, empty values;
- personal data that should not be in training data.
- For each issue, give the evidence rows, the count in the sample, the likely effect on the model, and a concrete fix: relabel with a guideline change, deduplicate by exact hash or by near-duplicate detection, split by group, drop or mask a leaking field, collect or reweight under-represented slices, or scrub personal data.
- Propose specific wording changes to the labelling guideline for every ambiguity you found.
- List the checks to run on the full dataset, such as cross-validated predictions to surface likely mislabels, near-duplicate detection across splits, and label distribution by source and by time.
- Refer to rows by id, or by row number if there is no id. Do not copy personal data into the report.
- Report counts as "n of N in the sample". Do not extrapolate a prevalence to the full dataset without saying it is an estimate from a sample of that size.
- If the sample is too small or clearly not random, say what it can and cannot show.
- Suggest a relabel only when you can say why. Mark your confidence as high, medium or low.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Summary
The three issues that matter most, one line each.
Findings
Table: issue | evidence rows | count in sample | effect on the model | fix.
Suspected mislabels
Table: row | current label | suggested label | reason | confidence.
Guideline changes
Bullets with the proposed wording.
Checks on the full dataset
Numbered, each with what it detects.
2 required values still a placeholder; the assistant will ask for them.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Data scientist, Data engineer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install review-training-data --target claude-codenpx skills add hermes-hq/hodios-dist --skill review-training-data -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringPlan a machine-learning experiment
Plans a machine-learning experiment before any training code exists: framing, baselines, leak-proof splits, metrics, ablations and a stop rule. Use when starting a new model or modelling spike.
plan-ml-experimentPlan a fine-tuning project
Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.
plan-fine-tuningMachine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
ml-engineerBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extractionChoose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
choose-ml-approach