Extract tables from a PDF
Extracts tables from PDF or scanned text into clean CSV, Markdown or JSON, keeps values exactly as printed, flags likely OCR errors and validates totals against the source. Use before analysis.
Text copied from PDFs loses table structure: columns run together, multi-line cells split into separate rows, headers span several lines, tables continue across pages with repeated headers, and scanned pages add OCR errors such as O for 0, l for 1 or a dropped decimal point. Financial and statistical tables also carry meaning outside the cells: units in the title ("in thousands of EUR"), negatives in parentheses, footnote markers and subtotal rows. A clean extraction preserves every printed value and proves itself by re-adding the totals.
Extract every table from this source as :
- Find each table and give it a name from its title or caption, with the page if shown.
- Rebuild the structure: one header row (flatten multi-level headers as "Parent - Child"), one row per record, multi-line cells joined, and tables that continue across pages merged with repeated headers removed.
- Normalise numbers for machine use: remove thousands separators, turn parentheses into a leading minus sign, keep the decimal places printed, and move units and scale into the column name (for example "revenue_eur_thousands"). Keep "-", "n/a" and blank cells as empty values and list them.
- Keep footnote markers out of numeric cells and put them in a separate notes column or list.
- Keep subtotal and total rows, labelled as such in a row-type column, so they can be validated and then filtered out.
- Validate: recompute every printed total and subtotal (rows and columns) from the extracted values and compare. Report each match and each mismatch with the difference.
- Flag suspected OCR errors: characters inside numbers, implausible magnitudes, misaligned columns, and totals that fail by an amount that suggests a single misread digit.
- Never change a printed value to make a total work. Report the mismatch and the likely culprit instead.
- Never fill an empty or unreadable cell with a guess; leave it empty and list it.
- Output each table in a separate fenced block. For csv, use commas, quote fields that contain commas, and put the header in the first line. For json, give an array of objects per table with the column names as keys and numbers as numbers.
- If the source has no recognisable table, or the text is too garbled to rebuild columns reliably, say so and ask for a better copy (for example exported text, a higher-resolution scan or the page image) rather than producing a doubtful table.
Tables found
For each table: its name, page, row and column count, then the table in a fenced block.
Validation
A table: table | total checked | printed | recomputed | result (match / mismatch and difference).
Issues to check
Bullets: empty cells, suspected OCR errors and structural guesses, each with its location.
1 required value still a placeholder; the assistant will ask for it.
details
- kind
- Prompt: a task you run by name to get one finished thing back
- domain
- Data analysis
- category
- Spreadsheets
- level
- Beginner
- made for
- Data analyst, Financial analyst, Researcher / scientist, Operations
- risk
- read-only
- version
- v1.0.0 · experimental
- reviewed
- 2026-10-02
- works in
- Claude Code, Codex, Cursor, GitHub Copilot, Gemini CLI, Antigravity, OpenCode, Windsurf, Zed, Continue, AGENTS.md, ChatGPT, claude.ai
use in
npx @hermes-hq/hodios install extract-tables-from-pdf --target claude-codenpx skills add hermes-hq/hodios-dist --skill extract-tables-from-pdf -a claude-codeclaude plugin marketplace add hermes-hq/hodios-distclaude plugin install hodios-data-analysis@hodiosThe plugin brings every entry in this domain at once.
pairs well with
All of SpreadsheetsReconcile two datasets
Reconciles two datasets that should agree, such as bank versus ledger or CRM versus billing, by matching records, listing mismatches and explaining likely causes. Use for month-end checks.
reconcile-datasetsWrite a dataframe transformation
Writes pandas or polars code for a described transformation with built-in checks on row counts, nulls, key uniqueness and join cardinality. Use when reshaping, joining or aggregating data.
write-dataframe-transformationRun a what-if analysis
Builds a scenario and sensitivity analysis for a decision (best, base and worst cases, a tornado chart and breakevens) as a spreadsheet layout with exact formulas. Use before committing to a plan.
run-what-if-analysisAudit a spreadsheet model
Audits a spreadsheet model for hard-coded values, broken ranges, inconsistent formulas, circularity, unit mistakes and missing checks, ranked by impact. Use before relying on someone else's sheet.
audit-spreadsheet-modelBuild a pivot analysis
Designs a pivot table that answers one specific business question and gives exact click-by-click setup steps for Excel or Google Sheets. Use when you have a flat table and a question about it.
build-pivot-analysisBuild a spreadsheet dashboard
Builds a working dashboard inside Excel or Google Sheets with a data tab, summary formulas, charts, slicers or dropdowns and refresh steps. Use when the team lives in spreadsheets, not a BI tool.
build-sheets-dashboard