hermes

Design a data pipeline

Designs a batch or streaming data pipeline sized to stated volumes, covering sources, schedule, idempotency, late data, backfills and monitoring. Use before building or replacing a pipeline.

context

Pipelines rarely fail on the happy path. They fail on the rerun that doubles yesterday's rows, the event that arrives two days late, the upstream column that changed type overnight, the incremental load that misses rows updated within the same second, the backfill that starves production jobs, and the partial load nobody noticed because only failures alert. Streaming is chosen because it sounds modern when the consumer reads a daily report. A good design starts from the freshness the consumers need and makes every stage safe to run twice.

task

Design a pipeline for: Only if [VOLUME] is given:

Volume: Only if [STACK] is given:

Stack:

  1. Pin down requirements: each source (type, how changes can be captured, rate limits), each destination, the consumers and their freshness need, delivery semantics (exactly-once effect, or at-least-once with deduplication), retention, and personal data handling. If freshness or volume is missing and would change the design, ask; otherwise state the assumption.
  2. Choose batch, micro-batch or streaming, justified by the freshness need and volume rather than preference. Size it: events or rows per second at peak, bytes per day, growth over two years, and the partitioning scheme that follows.
  3. Ingestion: change data capture, incremental extraction by a cursor column, or full snapshots. For cursor-based extraction, handle ties on the cursor value, clock skew and deletes that the cursor cannot see.
  4. Idempotency: make every stage safe to rerun by overwriting deterministic partitions or merging on keys, with deduplication keys and a run identifier recorded on output rows.
  5. Late and out-of-order data: event time versus processing time, the watermark or lookback window, and how corrections reach downstream tables.
  6. Schema evolution: the contract with each producer, what happens on a breaking change (fail, quarantine, or dead-letter), and who is told.
  7. Orchestration: the dependency graph, schedule, retries with backoff, timeouts and SLAs.
  8. Backfills: parameterised by date range, throttled, isolated from scheduled runs, and validated afterwards.
  9. Monitoring: freshness, volume, schema, null rates, consumer lag, rejected records and cost, each with a threshold, an owner, and whether it blocks publishing.
  10. List failure modes: what breaks, how it is detected, and how to recover.
constraints
  • Use the given stack. If none is given, use the fewest components that meet the requirements, and name alternatives only as examples.
  • Show the sizing arithmetic, and label numbers you supplied as assumptions.
  • Do not add streaming, a lakehouse, or a message bus unless a stated requirement needs it.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Summary

One paragraph, then a Mermaid or ASCII diagram of the flow.

Requirements and assumptions

Bullets, with assumptions marked.

Architecture

Stage by stage: what it does, the technology, and the schedule or trigger.

Idempotency and late data

How reruns and late events are handled at each stage.

Backfills

The procedure and its safeguards.

Monitoring and alerts

Table: signal | threshold | owner | blocks publishing (yes or no).

Failure modes

Table: failure | detection | recovery.

Sizing and cost

The arithmetic and the main cost drivers.

Open questions

Only those whose answers would change the design.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Data engineering
level
Expert
made for
Data engineer, Backend engineer, Software architect
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install design-data-pipeline --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill design-data-pipeline -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptData engineering

Write data-quality checks for a table

Writes data-quality checks for a table (freshness, volume, schema, validity, uniqueness, referential integrity, distribution) with severities, thresholds and owners. Use when a table feeds decisions.

write-data-quality-checks
PromptData engineering

Design a star schema

Designs a dimensional model from the questions analysts need answered: business processes, grain, facts, dimensions, slowly changing dimension types and DDL. Use when building a warehouse layer.

design-star-schema
PersonaData engineering

Data engineer

Acts as a data engineer who designs for idempotency, backfills and observability, treats schemas as contracts with their consumers, and asks who depends on each table before changing it.

data-engineer
PromptData engineering

Design a relational database schema

Designs a relational schema from requirements and access patterns, with keys, constraints, types, indexes and DDL. Use when starting a new service or feature that stores data.

design-database-schema
PromptData engineering

Generate realistic seed data

Generates realistic, referentially consistent fixture data for a database schema, with labelled edge cases and no real personal data. Use for local development, demos and integration tests.

generate-realistic-seed-data
PromptData engineering

Plan a zero-downtime schema change

Turns current table DDL and a desired change into expand and contract steps with lock-safe SQL, app changes, backfill, verification and rollback. Use before altering a live table.

plan-zero-downtime-schema-change