Research Paper to Markdown — Convert Academic PDFs
Academic papers are the worst-case PDF: two columns, dense citations, equations rendered as glyphs, figures with captions that wander. Generic PDF extractors mangle reading order. Our converter handles arXiv preprints, journal articles, and conference proceedings as a special case — column-aware, citation-aware, equation-aware.
What makes academic PDFs hard
Three problems compound. First, two-column layout: a naïve top-to-bottom extractor reads across columns and produces gibberish. Second, citations: in-text refs like [12] need to stay attached to their sentence, not float to a footnote. Third, math: rendered equations are positioned glyphs, not text — getting LaTeX back requires recognising the equation regions.
Our converter detects multi-column layouts and reads them in correct order, preserves in-text citations as inline references, converts displayed equations to LaTeX ($$...$$) and inline equations to $...$, and keeps figure captions with their figure numbers so you can find them later.
Reading order on column breaks
The converter analyses block bounding boxes per page, identifies column geometry, and emits text in the order a human reader would follow. Footnotes are collected at the end of the section that referenced them; references appear as a bibliography section at the end. The result reads like the paper, not like the file.
Before / After
Before (PDF):
[Two-column PDF]
Introduction Recent advances in transfer\nlearning [12] have demonstrated that\nlarge pretrained models... 12. Vaswani et al., 2017\n13. Devlin et al., 2018\n14. Brown et al., 2020
After (Markdown):
## Introduction
Recent advances in transfer learning [12] have demonstrated that large pretrained models...
## References
[12] Vaswani et al., 2017.
[13] Devlin et al., 2018.
[14] Brown et al., 2020.
Frequently asked questions
Does this preserve citations in academic papers?
Yes — in-text citations like [12] stay attached to their sentence, and the bibliography is emitted as a final ## References section. The numbering matches the original, so cross-references continue to work.
How are equations from academic PDFs handled?
Display equations are converted to LaTeX inside $$...$$; inline equations to $...$. The output renders correctly in any Markdown viewer with MathJax or KaTeX (Obsidian, MkDocs Material, GitHub readme rendering, etc.).
Can I convert arXiv papers directly?
Yes — download the PDF from arXiv (the version-pinned URL works best) and convert. arXiv's LaTeX-derived PDFs are particularly clean, so conversion fidelity is high. For batch academic ingestion, we recommend keeping the arXiv ID in your front matter for traceability.
Does this work on journal PDFs with dense layouts?
Yes for most major publishers (Elsevier, Springer, IEEE, ACM, Nature, Wiley). Layout heuristics handle their column geometries. Edge cases (legal-style margin notes, century-old typesetting) may need manual review of the converted Markdown.
Are figure captions and tables preserved?
Captions are kept attached to a placeholder for the figure (so you can re-insert images by hand if needed) and labelled with their original "Figure 1", "Table 2" identifiers. Tables are converted to GFM where layout permits; complex multi-row headers may need post-edit.