Design a data pipeline
Designs a batch or streaming data pipeline sized to stated volumes, covering sources, schedule, idempotency, late data, backfills and monitoring. Use before building or replacing a pipeline.
Pipelines rarely fail on the happy path. They fail on the rerun that doubles yesterday's rows, the event that arrives two days late, the upstream column that changed type overnight, the incremental load that misses rows updated within the same second, the backfill that starves production jobs, and the partial load nobody noticed because only failures alert. Streaming is chosen because it sounds modern when the consumer reads a daily report. A good design starts from the freshness the consumers need and makes every stage safe to run twice.
Design a pipeline for: Only if [VOLUME] is given:
Volume: Only if [STACK] is given:
Stack:
- Pin down requirements: each source (type, how changes can be captured, rate limits), each destination, the consumers and their freshness need, delivery semantics (exactly-once effect, or at-least-once with deduplication), retention, and personal data handling. If freshness or volume is missing and would change the design, ask; otherwise state the assumption.
- Choose batch, micro-batch or streaming, justified by the freshness need and volume rather than preference. Size it: events or rows per second at peak, bytes per day, growth over two years, and the partitioning scheme that follows.
- Ingestion: change data capture, incremental extraction by a cursor column, or full snapshots. For cursor-based extraction, handle ties on the cursor value, clock skew and deletes that the cursor cannot see.
- Idempotency: make every stage safe to rerun by overwriting deterministic partitions or merging on keys, with deduplication keys and a run identifier recorded on output rows.
- Late and out-of-order data: event time versus processing time, the watermark or lookback window, and how corrections reach downstream tables.
- Schema evolution: the contract with each producer, what happens on a breaking change (fail, quarantine, or dead-letter), and who is told.
- Orchestration: the dependency graph, schedule, retries with backoff, timeouts and SLAs.
- Backfills: parameterised by date range, throttled, isolated from scheduled runs, and validated afterwards.
- Monitoring: freshness, volume, schema, null rates, consumer lag, rejected records and cost, each with a threshold, an owner, and whether it blocks publishing.
- List failure modes: what breaks, how it is detected, and how to recover.
- Use the given stack. If none is given, use the fewest components that meet the requirements, and name alternatives only as examples.
- Show the sizing arithmetic, and label numbers you supplied as assumptions.
- Do not add streaming, a lakehouse, or a message bus unless a stated requirement needs it.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Summary
One paragraph, then a Mermaid or ASCII diagram of the flow.
Requirements and assumptions
Bullets, with assumptions marked.
Architecture
Stage by stage: what it does, the technology, and the schedule or trigger.
Idempotency and late data
How reruns and late events are handled at each stage.
Backfills
The procedure and its safeguards.
Monitoring and alerts
Table: signal | threshold | owner | blocks publishing (yes or no).
Failure modes
Table: failure | detection | recovery.
Sizing and cost
The arithmetic and the main cost drivers.
Open questions
Only those whose answers would change the design.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Data engineering
- level
- Expert
- made for
- Data engineer, Backend engineer, Software architect
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install design-data-pipeline --target claude-codenpx skills add hermes-hq/hodios-dist --skill design-data-pipeline -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of Data engineeringWrite data-quality checks for a table
Writes data-quality checks for a table (freshness, volume, schema, validity, uniqueness, referential integrity, distribution) with severities, thresholds and owners. Use when a table feeds decisions.
write-data-quality-checksDesign a star schema
Designs a dimensional model from the questions analysts need answered: business processes, grain, facts, dimensions, slowly changing dimension types and DDL. Use when building a warehouse layer.
design-star-schemaData engineer
Acts as a data engineer who designs for idempotency, backfills and observability, treats schemas as contracts with their consumers, and asks who depends on each table before changing it.
data-engineerDesign a relational database schema
Designs a relational schema from requirements and access patterns, with keys, constraints, types, indexes and DDL. Use when starting a new service or feature that stores data.
design-database-schemaGenerate realistic seed data
Generates realistic, referentially consistent fixture data for a database schema, with labelled edge cases and no real personal data. Use for local development, demos and integration tests.
generate-realistic-seed-dataPlan a zero-downtime schema change
Turns current table DDL and a desired change into expand and contract steps with lock-safe SQL, app changes, backfill, verification and rollback. Use before altering a live table.
plan-zero-downtime-schema-change