You don't want the whole PDF — you want the tables in it. Bank statements, financial reports, invoice line items, scientific data tables. We detect every table in the PDF and emit each as its own CSV file, ready to drop into a spreadsheet or pandas DataFrame.
Multi-table PDFs
Many PDFs contain multiple tables — a financial report might have ten. Conversion produces one CSV per detected table, named by sequence (page-3-table-1.csv) or by section heading when one is nearby. The output is a ZIP of CSVs plus a manifest noting where each table came from in the original PDF.
Bank statements and invoices
Two of the most common requests. Bank statements come out as one CSV per statement period with date, description, debit, credit, balance columns. Invoice line items come out as one CSV per invoice with description, quantity, unit price, line total columns. Headers and totals are kept on separate rows so you can drop them or aggregate as you prefer.
Frequently asked questions
How are multi-table PDFs handled?
Each detected table is emitted as its own CSV file, named by page and table index. The output is a ZIP archive with all CSVs plus a manifest mapping each CSV back to its location in the source PDF.
Will column headers be preserved?
Yes — when the table has clearly-distinguished header rows (typically the first row, often bold or with a different background), they become the first row of the CSV. Multi-row headers are flattened with / as a separator.
Can I extract tables from scanned PDFs?
Yes — OCR runs first, then table detection. Quality depends on the scan; cleanly-bordered tables in 200+ DPI scans usually come out cleanly. Borderless tables in scans are best-effort.
Best for bank statements and invoices?
Yes — both are tabular by nature and convert cleanly to CSV. Bank statements export as one CSV per period with date/description/amount/balance columns. Invoice line items export as one CSV per invoice with description/qty/price/total columns.
How does this compare to dedicated table extractors?
Dedicated tools (Tabula, Camelot, Excalibur) require setup and per-PDF configuration to get good results. Our converter handles most common layouts out of the box; for unusual or highly customised tables, those dedicated tools may produce cleaner results with manual tuning.