Write a data dictionary
Writes a data dictionary for database tables with each column's meaning, units, nullability, allowed values, owner and lineage, and flags every column it cannot infer. Use when documenting a schema.
A data dictionary is only trusted if it never guesses silently. The expensive mistakes come from the columns that look obvious: amount stored in cents and read as currency units, created_at in local time read as UTC, a status code 3 nobody can decode, a nullable column whose nulls mean "not applicable" in one era and "unknown" in another. The value of the dictionary is as much in naming what is not known, and whom to ask, as in describing what is.
Write a data dictionary for: Only if [SAMPLE_ROWS] is given:
Sample rows:
- For each table, state the grain ("one row per …"), the primary key, and how rows appear to change (append-only, updated in place, soft-deleted), if the evidence shows it.
- For each column, record:
- meaning, in one plain sentence;
- unit or format (currency and minor units, time zone, ID format, encoding);
- nullability, declared and observed in the sample, and what a null means;
- allowed values or range, from constraints or observed in the sample;
- an example value (masked if sensitive);
- personal data classification: none, personal or sensitive;
- lineage: the foreign key it references, or what it is derived from;
- owner, or "TBD";
- confidence: declared (from constraints or comments), inferred (from the name or sample), or unknown.
- Where you cannot infer the meaning or unit with confidence, write "Cannot infer" and add a precise question for the owner, for example "Is orders.amount in cents or in currency units? Sample values 1999 and 250 suggest cents."
- Flag inconsistencies: the same concept named differently across tables, mixed units, columns that look unused or always null in the sample, and codes without a lookup table.
- Never present an inference as a fact. Every inferred entry is marked as inferred.
- Do not copy personal data from the sample into the dictionary. Mask example values.
- Keep each meaning to one sentence. Put detail in the questions, not in the table.
For each table, a heading ## <table name>, a one-line summary (grain, key, change pattern), then a table: Column | Type | Meaning | Unit or format | Nullable (declared/observed) | Allowed values | Example | PII | Lineage | Owner | Confidence.
Finish with ## Questions for owners: a numbered list grouped by table, each question answerable in one line.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Data engineering
- level
- Intermediate
- made for
- Data engineer, Data analyst, Backend engineer, Technical writer
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install write-data-dictionary --target claude-codenpx skills add hermes-hq/hodios-dist --skill write-data-dictionary -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Data engineeringWrite data-quality checks for a table
Writes data-quality checks for a table (freshness, volume, schema, validity, uniqueness, referential integrity, distribution) with severities, thresholds and owners. Use when a table feeds decisions.
write-data-quality-checksData engineer
Acts as a data engineer who designs for idempotency, backfills and observability, treats schemas as contracts with their consumers, and asks who depends on each table before changing it.
data-engineerDesign a data pipeline
Designs a batch or streaming data pipeline sized to stated volumes, covering sources, schedule, idempotency, late data, backfills and monitoring. Use before building or replacing a pipeline.
design-data-pipelineDesign a relational database schema
Designs a relational schema from requirements and access patterns, with keys, constraints, types, indexes and DDL. Use when starting a new service or feature that stores data.
design-database-schemaDesign a star schema
Designs a dimensional model from the questions analysts need answered: business processes, grain, facts, dimensions, slowly changing dimension types and DDL. Use when building a warehouse layer.
design-star-schemaGenerate realistic seed data
Generates realistic, referentially consistent fixture data for a database schema, with labelled edge cases and no real personal data. Use for local development, demos and integration tests.
generate-realistic-seed-data