OCR alone gives you raw text. OCR plus Markdown conversion gives you structure: headings, lists, tables, paragraphs — what the document originally was. We bundle both steps so the output is immediately usable, not just searchable.
OCR is necessary but not sufficient
Most OCR tools return a text file: walls of recognised characters with paragraph breaks if you're lucky. That's enough to pass through full-text search but useless for AI workflows that depend on structure. Our pipeline runs OCR first, then applies the same layout-recognition heuristics it uses on digital PDFs to recover headings, lists, and tables. The output is Markdown — not just OCR'd text.
What "OCR + structure" actually does
After character recognition, we run text-layout analysis on the recognised text positions: large-text regions become headings, indented bullet-like patterns become lists, gridded text becomes a table candidate. The result is the closest thing to "what the original Word/InDesign file would have looked like" that's achievable from a scan.
Frequently asked questions
How is OCR PDF to Markdown different from regular OCR?
Regular OCR gives you raw recognised text. We bundle layout analysis after recognition, so the output has Markdown structure — headings, lists, tables — rather than a plain text dump. That difference is what makes the result useful for AI workflows and Markdown-based wikis.
What languages does the OCR engine support?
English, French, German, Spanish, Italian, Portuguese, Dutch, Polish, Czech, plus most Latin-script languages with diacritics. Cyrillic (Russian, Ukrainian) and CJK (Chinese, Japanese, Korean) are supported with separate language packs that we auto-select per document.
Can I review OCR confidence scores?
On request, the API exposes per-region confidence scores so you can flag low-confidence text for manual review. The web UI marks suspect regions with a subtle highlight in the preview.
Are scanned tables converted to Markdown tables?
When the table has clear borders and consistent cell alignment, yes — GFM tables are emitted. Borderless tables are best-effort; you may need to spot-check alignment.
How long does OCR take for a 100-page scan?
Roughly 30–90 seconds depending on image quality and language complexity. The bottleneck is OCR itself; layout analysis adds <10% on top.