Multi-Column PDF to Markdown — Correct Reading Order
A two-column PDF read top-to-bottom is gibberish — half a sentence from column 1, then a jump back to the top of column 2. Our converter analyses column geometry on every page and emits text in the order you actually read it.
Why naïve extraction fails on columns
The simplest PDF text extractor walks glyphs in their stored order. That order is usually generation order — whatever the PDF generator emitted as it laid out the page. For multi-column documents that order is essentially random. Some generators emit each column top-to-bottom (the right answer), others interleave columns by Y-coordinate (disastrous), some sort by generation date (also disastrous).
Column-aware extraction
Our converter clusters text blocks by their X-coordinate centroids to detect column boundaries, then emits text in column order top-to-bottom. Section headings that span columns (full-width ## Heading) are recognised and placed in their proper position. Footnotes attached to a column flow with that column. Two-column papers, three-column magazines, and four-column brochures all pass through correctly.
Before / After
Before (PDF):
[Two columns read in stored order:]
The quick brown fox jumps over
fox jumps over the lazy dog
the lazy dog in the meadow
After (Markdown):
## Article Title
The quick brown fox jumps over the lazy dog in the meadow.
Frequently asked questions
How is reading order determined for multi-column PDFs?
We cluster text blocks by X-coordinate to detect column boundaries, then emit each column top-to-bottom. Section headings that span all columns are placed at the right position in the resulting flow.
Does this work on three-column or four-column layouts?
Yes — the column-detection algorithm is general. Three-column academic papers, four-column brochures, and even the rare five-column reference layouts all extract in correct order. Anything beyond five columns is rare enough that we recommend manual review.
What if columns have different widths?
Asymmetric layouts (a wide column with a narrower sidebar, common in magazines) are detected as separate columns of different widths. The wider column is emitted first, which matches the natural reading order of nearly all such layouts.
How are full-width section headings handled?
Recognised by their bounding box spanning the column boundaries, and emitted as ## headings between the column blocks at the position they appeared. The result reads exactly like the original page.
What about headers and footers in multi-column PDFs?
Repeating page-level headers and footers are detected and stripped, regardless of column count. They're considered furniture, not content.