Design an event-driven system
Designs an event-driven flow with event schemas, topics, partition keys, idempotent consumers, an outbox, retries, dead letters and replay. Use when moving synchronous calls onto a broker.
Moving a flow from synchronous calls to a broker trades one set of failure modes for another. Teams usually get the happy path right and then meet the hard parts in production: the database commit succeeds but the publish fails (or the reverse), a consumer processes the same message twice because delivery is at-least-once, events for the same order arrive out of order because the partition key was wrong, a poison message blocks a partition, a schema change breaks a consumer nobody knew about, and nobody can replay a week of events after a bug. A good design decides each of these explicitly, and also says plainly when a synchronous call is still the better choice for a step.
Design the event-driven version of this flow:
Only if [THROUGHPUT] is given: Expected throughput: Broker: Only if [CONSISTENCY_NEEDS] is given: Consistency needs:
- If the flow, the services involved or the consistency needs are too vague to decide ordering and delivery guarantees, ask up to five specific questions and stop. Otherwise continue, labelling every assumption.
- Map the flow: the steps, which service owns each, and for each step whether it should be an event (something that happened, owned by its producer), a command (a request for one specific service to act) or stay a synchronous call (when the caller needs the answer to proceed). Justify each choice in one line.
- Define the event catalogue. Name events in the past tense in domain language (OrderPlaced, PaymentCaptured). For each: producer, consumers, trigger, payload fields with types, and whether it carries the full state (event-carried state transfer) or only ids (notification). Every event has an envelope with event id, type, schema version, occurred-at time in UTC, producer, correlation id and causation id; prefer the CloudEvents attribute names unless the team already has a convention.
- Design the topology: topics, queues or streams; partition or ordering keys chosen from the entity whose events must stay in order; partition counts sized from the throughput with the arithmetic shown; retention; and consumer groups. State exactly which ordering is guaranteed (per key, never global) and what happens to it during retries and rebalances.
- Make publishing reliable: use a transactional outbox (or change data capture on the outbox table) so the state change and the event commit together; describe the relay, its ordering and how it avoids publishing duplicates where it can. Say why dual writes are unsafe here.
- Make consumers idempotent: assume at-least-once delivery, choose the deduplication strategy per consumer (a processed-message table keyed by event id written in the same transaction as the side effect, natural idempotency, or version checks), and handle out-of-order events with entity versions or by fetching current state.
- Define failure handling: retry policy with exponential backoff and jitter, which errors are retryable, retry topics or delayed redelivery versus blocking retries, a dead-letter destination per consumer with the original payload and error metadata, alerting, and the runbook for inspecting, fixing and redriving dead letters. For multi-step business transactions, design the saga (choreography or orchestration, with the choice justified) and the compensating actions.
- Plan replay and evolution: how a consumer rebuilds state from retained events or a snapshot, how to reprocess safely given idempotency, schema registry or contract checks, compatible-change rules (add optional fields; never rename or repurpose), and how a breaking change ships as a new event version alongside the old.
- List what to observe: consumer lag per group, end-to-end latency from occurred-at, dead-letter counts, outbox backlog, duplicate rate, and the alerts on each.
- If is "any", recommend a broker for this throughput, ordering and team and explain the deciding factors. Otherwise use the named broker's own concepts and limits, and say where a feature you rely on differs by broker.
- Do not introduce events where a synchronous call is simpler and the caller needs the result; say so instead.
- Never claim exactly-once delivery end to end. If the broker offers transactional or exactly-once features, state precisely what they cover and what still needs idempotent consumers.
- Do not invent broker limits, quotas or prices. When a number matters and you are not sure of it, say how to look it up.
- Keep business rules you were not given as marked assumptions.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
Summary
The design in at most 6 lines, including the broker and the delivery guarantee.
Flow
A Mermaid sequence or flowchart diagram, then a table: step, owner, event or command or sync call, why.
Event catalogue
Table: event, producer, consumers, partition key, payload fields, state or notification. Then one example event as JSON with its envelope.
Topology and ordering
Topics or queues with partitions, retention and consumer groups, and the sizing arithmetic.
Producers and the outbox
The outbox table, the relay and publish guarantees.
Consumers and idempotency
Per consumer: dedup strategy, ordering handling, side effects.
Failure handling
Retry policy, dead letters, redrive runbook and any saga with compensations.
Replay and evolution
Observability
Metrics and alerts as a table.
Assumptions and open questions
Numbered. Each says what it affects.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Software engineering
- category
- Architecture
- level
- Expert
- made for
- Backend engineer, Software architect, Tech lead / staff engineer, Data engineer
- risk
- read-only
- version
- v1.0.0 · incubating
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install design-event-driven-system --target claude-codenpx skills add hermes-hq/hodios-dist --skill design-event-driven-system -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-software-engineering@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of ArchitectureSoftware architect
Acts as a pragmatic software architect who designs from requirements and constraints, names trade-offs and failure modes, and keeps designs as simple as the problem allows.
software-architectBackend engineer
Acts as a backend engineer focused on correct data handling, clear API contracts, explicit failure modes and services that are easy to operate. Use as a builder or reviewer persona for server code.
backend-engineerWrite an architecture decision record
Writes an architecture decision record that states one decision, the forces behind it, the options weighed and the honest consequences. Use when a significant technical choice is made or proposed.
write-adrReview a system design
Reviews a design document or proposal for failure modes, scaling limits, data and consistency risks and operability gaps, and returns ranked findings. Use before a design review or before building.
review-system-designImplement a background job
Implements a background or scheduled job with idempotency, retries with backoff, dead-letter handling, timeouts, concurrency limits and observability. Use to move slow work off the request path.
implement-background-jobCompare design options
Compares two to four technical options against the criteria that matter, weighs reversibility and risk, and recommends one. Use when a team is stuck choosing between approaches or tools.
compare-design-options