Write a model card
Writes a model card with intended use, training data, metrics by slice, limitations and ethical considerations from training notes and eval results, flagging gaps. Use before releasing a model.
A model card tells someone deciding whether to use a model what it is for, what it was trained and tested on, where it works and where it fails. Readers include engineers integrating it, reviewers approving its release and people affected by its decisions. Weak cards read like marketing: one headline metric, no slices, limitations that are generic or invented, and no out-of-scope uses. A useful card states only what the evidence supports and says plainly what was never measured.
Write a model card from these notes: Only if [EVAL_RESULTS] is given:
Evaluation results:
- Fill each section from the evidence: model details (name, version, type, architecture or base model, date, owner, license), intended use and users, out-of-scope uses, training data (sources, size, time range, preprocessing, known gaps), evaluation data, metrics, limitations, ethical considerations, and recommendations for users.
- Derive out-of-scope uses from the evidence. For example, training data in one language makes other languages out of scope, and data from one period makes later periods unverified.
- Report metrics overall and by every slice available, with sample sizes and confidence intervals where they exist. Call out the largest gap between slices with its numbers.
- Where the notes say nothing, write "Not documented" and add a precise question to Gaps to fill naming who or what could answer it.
- Flag contradictions between the notes and the results, such as a claim of multilingual support with English-only evaluation.
- Never invent a number, dataset, license or limitation. Mark anything you inferred as an inference.
- Do not round or average away a disparity between slices.
- Write for a technical reader who is not on the team, in plain language, defining any metric name a reader may not know.
- Keep marketing language out ("state-of-the-art", "robust", "unbiased").
A Markdown model card with these headings, in order: Model details, Intended use, Out-of-scope uses, Training data, Evaluation data, Metrics, Limitations, Ethical considerations, Recommendations, Gaps to fill. Present metrics as a table: slice | metric | value | sample size. Gaps to fill is a numbered list of questions.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- AI and ML engineering
- level
- Intermediate
- made for
- ML / AI engineer, Data scientist, Researcher / scientist, Technical writer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install write-model-card --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-model-card -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of AI and ML engineeringMachine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
ml-engineerBuild an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
build-mcp-serverBuild an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
build-structured-extractionChoose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
choose-ml-approachDesign an LLM agent architecture
Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.
design-agent-architectureDesign a RAG pipeline
Designs a retrieval-augmented generation pipeline from a corpus and its real questions, covering chunking, hybrid retrieval, reranking, citations and evals. Use before building or rebuilding RAG.
design-rag-pipeline