hermes

Extract tables from a PDF

Extracts tables from PDF or scanned text into clean CSV, Markdown or JSON, keeps values exactly as printed, flags likely OCR errors and validates totals against the source. Use before analysis.

context

Text copied from PDFs loses table structure: columns run together, multi-line cells split into separate rows, headers span several lines, tables continue across pages with repeated headers, and scanned pages add OCR errors such as O for 0, l for 1 or a dropped decimal point. Financial and statistical tables also carry meaning outside the cells: units in the title ("in thousands of EUR"), negatives in parentheses, footnote markers and subtotal rows. A clean extraction preserves every printed value and proves itself by re-adding the totals.

task

Extract every table from this source as :

source

  1. Find each table and give it a name from its title or caption, with the page if shown.
  2. Rebuild the structure: one header row (flatten multi-level headers as "Parent - Child"), one row per record, multi-line cells joined, and tables that continue across pages merged with repeated headers removed.
  3. Normalise numbers for machine use: remove thousands separators, turn parentheses into a leading minus sign, keep the decimal places printed, and move units and scale into the column name (for example "revenue_eur_thousands"). Keep "-", "n/a" and blank cells as empty values and list them.
  4. Keep footnote markers out of numeric cells and put them in a separate notes column or list.
  5. Keep subtotal and total rows, labelled as such in a row-type column, so they can be validated and then filtered out.
  6. Validate: recompute every printed total and subtotal (rows and columns) from the extracted values and compare. Report each match and each mismatch with the difference.
  7. Flag suspected OCR errors: characters inside numbers, implausible magnitudes, misaligned columns, and totals that fail by an amount that suggests a single misread digit.
constraints
  • Never change a printed value to make a total work. Report the mismatch and the likely culprit instead.
  • Never fill an empty or unreadable cell with a guess; leave it empty and list it.
  • Output each table in a separate fenced block. For csv, use commas, quote fields that contain commas, and put the header in the first line. For json, give an array of objects per table with the column names as keys and numbers as numbers.
  • If the source has no recognisable table, or the text is too garbled to rebuild columns reliably, say so and ask for a better copy (for example exported text, a higher-resolution scan or the page image) rather than producing a doubtful table.
output format

Tables found

For each table: its name, page, row and column count, then the table in a fenced block.

Validation

A table: table | total checked | printed | recomputed | result (match / mismatch and difference).

Issues to check

Bullets: empty cells, suspected OCR errors and structural guesses, each with its location.

1 required value still a placeholder; the assistant will ask for it.

details

kind
Prompt: a task you run by name to get one finished thing back
domain
Data analysis
category
Spreadsheets
level
Beginner
made for
Data analyst, Financial analyst, Researcher / scientist, Operations
risk
read-only
version
v1.0.0 · experimental
reviewed
2026-10-02
works in
Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai

Edit on GitHubReport a problem

use in

Hodios CLI
npx @hermes-hq/hodios install extract-tables-from-pdf --target claude-code
Agent Skills
npx skills add hermes-hq/hodios-dist --skill extract-tables-from-pdf -a claude-code
Add the Hodios marketplace (once)
claude plugin marketplace add hermes-hq/hodios-dist
Install the data-analysis plugin
claude plugin install hodios-data-analysis@hodios

The plugin brings every entry in this domain at once.

pairs well with

All of Spreadsheets
PromptData exploration

Reconcile two datasets

Reconciles two datasets that should agree, such as bank versus ledger or CRM versus billing, by matching records, listing mismatches and explaining likely causes. Use for month-end checks.

reconcile-datasets
PromptData exploration

Write a dataframe transformation

Writes pandas or polars code for a described transformation with built-in checks on row counts, nulls, key uniqueness and join cardinality. Use when reshaping, joining or aggregating data.

write-dataframe-transformation
PromptSpreadsheets

Run a what-if analysis

Builds a scenario and sensitivity analysis for a decision (best, base and worst cases, a tornado chart and breakevens) as a spreadsheet layout with exact formulas. Use before committing to a plan.

run-what-if-analysis
PromptSpreadsheets

Audit a spreadsheet model

Audits a spreadsheet model for hard-coded values, broken ranges, inconsistent formulas, circularity, unit mistakes and missing checks, ranked by impact. Use before relying on someone else's sheet.

audit-spreadsheet-model
PromptSpreadsheets

Build a pivot analysis

Designs a pivot table that answers one specific business question and gives exact click-by-click setup steps for Excel or Google Sheets. Use when you have a flat table and a question about it.

build-pivot-analysis
PromptSpreadsheets

Build a spreadsheet dashboard

Builds a working dashboard inside Excel or Google Sheets with a data tab, summary formulas, charts, slicers or dropdowns and refresh steps. Use when the team lives in spreadsheets, not a BI tool.

build-sheets-dashboard