hermes

Generate realistic seed data

Generates realistic, referentially consistent fixture data for a database schema, with labelled edge cases and no real personal data. Use for local development, demos and integration tests.

context

Seed data is only useful if it loads and if it looks like production. Typical generated fixtures fail on the first foreign key, use "test1, test2" names that hide layout bugs, give every customer exactly one order, and leave out the rows that break code: the longest name, the null middle name, the order with no items, the timestamp on a daylight-saving boundary. Fixtures also leak real personal data when someone copies production rows. Good seed data obeys every constraint, has realistic skew and ordering, deliberately includes edge cases, and is fictitious by construction.

task

Generate seed data as for the schema below. Row counts: .

  1. Parse the schema. Order tables so every referenced row exists before it is referenced. Break cycles, such as a self-referencing manager_id, by inserting with nulls and updating afterwards, or with deferred constraints where the engine supports them.
  2. Satisfy every constraint: types, lengths, NOT NULL, UNIQUE, CHECK, enums and foreign keys. For sql, write in the dialect the DDL implies and say which one you assumed.
  3. Make it realistic:
  • skewed relationships (a few customers with many orders, most with one or two);
  • timestamps in a consistent order (created before updated, ordered before shipped) relative to a fixed anchor date you state;
  • derived values that agree (an order total equals the sum of its lines). For sql, insert the parent with a placeholder its constraints accept (such as 0) and set the value with an UPDATE from the children, instead of doing the arithmetic by hand; for csv and json, recheck each one before output;
  • varied, plausible, invented names and text in several locales.
  1. Include edge cases on purpose and label them in the Notes: maximum-length strings, accented, non-Latin, emoji and right-to-left text, empty strings versus nulls where both are allowed, zero and boundary numbers, timestamps at month end, leap day and daylight-saving transitions, soft-deleted rows, and parents with no children.
  2. Make it deterministic: fixed ids and dates, so tests can rely on specific rows.
  3. Output: for sql, INSERT statements in dependency order inside one transaction; for csv, one block per table with a header row; for json, one object keyed by table name.
  4. If the requested volume is too large to list usefully (more than a few hundred rows in total), write a small hand-made set with the edge cases plus a deterministic, seeded generator for the bulk, and say why.
constraints
  • No real people, real companies' customer data, real addresses or working contact details. Use reserved example domains (example.com, example.org, example.net), fictional phone ranges such as 555-0100 to 555-0199 in North America, documentation IP ranges (192.0.2.0/24, 198.51.100.0/24, 203.0.113.0/24), and payment card numbers only from published test ranges.
  • For national identifiers and similar sensitive fields, use values that are structurally invalid or from documented test ranges, and say so.
  • If a column's meaning is unclear (a polymorphic type column, a JSON payload with no schema), ask or state the assumption.
output format

Notes

The insertion order, the anchor date, how cycles were broken, and a list of edge cases with the rows that carry them.

Data

The data in , in fenced blocks.

Constraint check

One line per constraint, saying how the data satisfies it.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Data engineering
level
Intermediate
made for
Software engineer, Backend engineer, QA / test engineer, Data engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install generate-realistic-seed-data --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill generate-realistic-seed-data -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

more in data engineering

All of Data engineering
PersonaData engineering

Data engineer

Acts as a data engineer who designs for idempotency, backfills and observability, treats schemas as contracts with their consumers, and asks who depends on each table before changing it.

data-engineer
PromptData engineering

Design a data pipeline

Designs a batch or streaming data pipeline sized to stated volumes, covering sources, schedule, idempotency, late data, backfills and monitoring. Use before building or replacing a pipeline.

design-data-pipeline
PromptData engineering

Design a relational database schema

Designs a relational schema from requirements and access patterns, with keys, constraints, types, indexes and DDL. Use when starting a new service or feature that stores data.

design-database-schema
PromptData engineering

Design a star schema

Designs a dimensional model from the questions analysts need answered: business processes, grain, facts, dimensions, slowly changing dimension types and DDL. Use when building a warehouse layer.

design-star-schema
PromptData engineering

Plan a zero-downtime schema change

Turns current table DDL and a desired change into expand and contract steps with lock-safe SQL, app changes, backfill, verification and rollback. Use before altering a live table.

plan-zero-downtime-schema-change
PromptData engineering

Review database indexes against the workload

Reviews a database's indexes against its real query workload to find missing, unused, duplicate and bloated indexes, with DDL and the write cost of each change. Use for periodic index hygiene.

review-database-indexes