Scanned PDF to Markdown — OCR + Conversion in One Step
A "scanned" PDF is just images of pages — there's no text underneath. Most PDF tools stare at it blankly. Our converter detects when a PDF has no text layer, runs OCR automatically, and emits Markdown the same way as for digital sources. One upload, one download, no separate OCR step.
How we detect a scanned PDF
The first heuristic is text density: if the PDF's text layer covers less than 10% of typical page area, we treat it as scanned. The second is glyph quality: fonts that look like the PDF was OCR'd badly already (mojibake, garbled spacing) get re-OCR'd. Either way, the user sees the same flow — drop the file, get Markdown back.
What OCR can and can't recover
Cleanly-scanned typed text at 300+ DPI: near-perfect, often above 99% accuracy. Lower-resolution scans (faxed contracts, photos taken on a phone): 90–98% accuracy with occasional confused pairs (l/1, O/0). Handwriting: variable — neat block printing works, cursive notes are a coin-flip. Tables in scanned form: rows and columns recovered, but cell alignment is best-effort.
Before / After
Before (PDF):
[Scanned image of a page]
(no extractable text — all glyphs are pixels)
After (Markdown):
# Service Agreement
This Agreement is entered into as of January 15, 2026, by and between Acme Corp ("Provider") and Beta LLC ("Client").
## 1. Scope of Services
Provider shall deliver...
Frequently asked questions
How does OCR handle scanned PDFs in this tool?
Automatically. We detect the absence of an extractable text layer (or a corrupted one) and run OCR before conversion. The output is regular Markdown — you don't have to know it was a scan.
What's the accuracy on low-quality scans?
For 200+ DPI typed scans, typically 95–99%. For 100–150 DPI scans (older fax-quality material), 85–95%. Below that, results degrade quickly. Phone-photographed pages depend heavily on lighting and angle but usually fall in the 90–98% range.
Can I OCR handwritten PDFs?
Block printing and clean cursive: usable, with manual review for confused characters. Doctor-style or rapid notes: not reliable enough to trust without careful proofreading. The OCR engine flags low-confidence regions in the output for you to inspect.
Does OCR slow down conversion significantly?
Scanned-PDF conversion takes ~3–10× longer than digital-PDF conversion of the same page count, because OCR is the dominant cost. For a 50-page scan, expect 15–60 seconds depending on quality.
Are scanned tables converted correctly?
Row and column structure is recovered for most clearly-bordered tables. Borderless tables, merged cells, or tables with rotated text are best-effort and often need manual cleanup. The converter outputs GFM tables when confidence is high and falls back to plain text rows otherwise.