hermes

Write data-quality checks for a table

Writes data-quality checks for a table (freshness, volume, schema, validity, uniqueness, referential integrity, distribution) with severities, thresholds and owners. Use when a table feeds decisions.

context

Most bad data is not a failed job. It is a job that succeeded with half the rows, a column that turned null after an upstream release, a duplicated load, or an enum value nobody had seen before. Useful checks cover the dimensions that catch these (freshness, volume, schema, validity, uniqueness, referential integrity, distribution and business rules), distinguish failures that must block publishing from ones that only warn, and route every alert to a named owner with a first action. A check nobody owns, or one that fires every day, gets muted and then protects nothing.

task

Write data-quality checks in for this table: Only if [SAMPLE_ROWS] is given:

Sample or statistics:

  1. State the grain ("one row per …"), the key, the load cadence and the consumers. If the grain or cadence is unclear, ask, or state the assumption.
  2. Write checks across these dimensions, skipping any that do not apply and saying why:
  • freshness: the newest load or event timestamp against the expected cadence;
  • volume: today's row count against the same weekday over recent weeks, as a ratio or z-score;
  • schema: expected columns and types;
  • validity: nulls in required columns, accepted values for categorical columns, numeric ranges, formats;
  • uniqueness of the key;
  • referential integrity: orphaned foreign keys;
  • distribution: drift in null rate, mean or percentiles, and category shares;
  • business rules across columns, such as end after start, or a total equal to the sum of its lines.
  1. Give each check a severity: block (stop downstream publishing) or warn. Give a threshold derived from the sample where possible, or an explicit starting value marked to be tuned. Name an owner role or a placeholder, and give the first action on failure.
  2. Implement the checks in :
  • sql: one query per check that returns failing rows or a single failing metric, so zero rows means pass;
  • dbt: generic tests in properties YAML plus singular tests, naming any package a test needs;
  • great-expectations: an expectation suite using the GX Core 1.x API (say which version you assumed);
  • soda: SodaCL checks in YAML.
  1. Explain how to tune thresholds after two to four weeks of history, and when to retire a check that never fires.
constraints
  • Do not invent columns. Checks must reference only columns in the table definition.
  • Avoid checks that will alert on normal variation. Weekly seasonality and month-end peaks belong in the threshold.
  • Keep each check independent, so one failure does not hide another.
  • Separate what you verified from what you inferred. Mark inferences as such.
  • When you do not know, say "I don't know" once and state what would settle it.
output format

Table grain and assumptions

Grain, key, cadence, consumers, and assumptions.

Checks

Table: check | dimension | severity | threshold | owner | first action on failure.

Implementation

The code for in fenced blocks, one per file.

Tuning plan

How and when to adjust thresholds.

Gaps

What these checks cannot catch, and what would.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Software engineering
category
Data engineering
level
Intermediate
made for
Data engineer, Data analyst, Site reliability engineer
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install write-data-quality-checks --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill write-data-quality-checks -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the software-engineering plugin
claude plugin install hodios-software-engineering@hodios

The plugin brings every entry in this domain at once.

PromptData engineering

Write a dbt model

Writes a dbt model from business logic, with declared sources, a stated grain, unique, not_null and relationships tests, column docs and a safe incremental strategy. Use when adding a dbt model.

write-dbt-model
PromptData engineering

Design a data pipeline

Designs a batch or streaming data pipeline sized to stated volumes, covering sources, schedule, idempotency, late data, backfills and monitoring. Use before building or replacing a pipeline.

design-data-pipeline
PromptData engineering

Write a data dictionary

Writes a data dictionary for database tables with each column's meaning, units, nullability, allowed values, owner and lineage, and flags every column it cannot infer. Use when documenting a schema.

write-data-dictionary
PersonaData engineering

Data engineer

Acts as a data engineer who designs for idempotency, backfills and observability, treats schemas as contracts with their consumers, and asks who depends on each table before changing it.

data-engineer
PromptData engineering

Design a relational database schema

Designs a relational schema from requirements and access patterns, with keys, constraints, types, indexes and DDL. Use when starting a new service or feature that stores data.

design-database-schema
PromptData engineering

Design a star schema

Designs a dimensional model from the questions analysts need answered: business processes, grain, facts, dimensions, slowly changing dimension types and DDL. Use when building a warehouse layer.

design-star-schema