hermes

Review a training dataset sample

Audits a sample of a labelled dataset for label noise, leakage, duplicates, class imbalance and representation gaps, and gives a concrete fix for each problem. Use before training or fine-tuning.

context

A model cannot be more consistent than its labels. Most dataset problems are systematic: a guideline that two labellers read differently, a source whose rows are all one class, a field that leaks the label, or thousands of near-identical rows that inflate test scores. Reading a sample row by row finds these problems far more cheaply than training a model and wondering why it plateaus. The aim is to find the patterns behind individual errors, not to relabel the sample.

task

Audit this sample for the task below.

Task:

Sample:

  1. Identify what one row represents, which column is the label, and the label set. If the label column or a label's meaning is unclear, ask before auditing.
  2. Read every row and check for:
  • label noise: rows whose label contradicts their content. Separate clear errors from ambiguous rows that reveal a guideline gap;
  • inconsistency: near-identical rows with different labels;
  • duplicates and near-duplicates, and across splits if there is a split column;
  • leakage: fields or text that give away the label (label words in the text, status tags, identifiers, timestamps recorded after the outcome, boilerplate unique to one source);
  • class balance: counts per label in the sample;
  • representation gaps: languages, lengths, sources, time periods or user groups that are missing or rare, and the edge cases the task implies but the sample lacks;
  • formatting defects: truncation, encoding errors, HTML or template residue, empty values;
  • personal data that should not be in training data.
  1. For each issue, give the evidence rows, the count in the sample, the likely effect on the model, and a concrete fix: relabel with a guideline change, deduplicate by exact hash or by near-duplicate detection, split by group, drop or mask a leaking field, collect or reweight under-represented slices, or scrub personal data.
  2. Propose specific wording changes to the labelling guideline for every ambiguity you found.
  3. List the checks to run on the full dataset, such as cross-validated predictions to surface likely mislabels, near-duplicate detection across splits, and label distribution by source and by time.
constraints
  • Refer to rows by id, or by row number if there is no id. Do not copy personal data into the report.
  • Report counts as "n of N in the sample". Do not extrapolate a prevalence to the full dataset without saying it is an estimate from a sample of that size.
  • If the sample is too small or clearly not random, say what it can and cannot show.
  • Suggest a relabel only when you can say why. Mark your confidence as high, medium or low.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Summary

The three issues that matter most, one line each.

Findings

Table: issue | evidence rows | count in sample | effect on the model | fix.

Suspected mislabels

Table: row | current label | suggested label | reason | confidence.

Guideline changes

Bullets with the proposed wording.

Checks on the full dataset

Numbered, each with what it detects.

2 required values still a placeholder; the assistant will ask for them.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
AI and ML engineering
level
Intermediate
made for
ML / AI engineer, Data scientist, Data engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install review-training-data --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill review-training-data -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptAI and ML engineering

Plan a machine-learning experiment

Plans a machine-learning experiment before any training code exists: framing, baselines, leak-proof splits, metrics, ablations and a stop rule. Use when starting a new model or modelling spike.

plan-ml-experiment
PromptAI and ML engineering

Plan a fine-tuning project

Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.

plan-fine-tuning
PersonaAI and ML engineering

Machine-learning engineer

Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.

ml-engineer
PromptAI and ML engineering

Build an MCP server

Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.

build-mcp-server
PromptAI and ML engineering

Build an LLM structured extraction step

Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.

build-structured-extraction
PromptAI and ML engineering

Choose between rules, ML and an LLM

Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.

choose-ml-approach