PDF to Markdown for LLMs: The Universal AI Input Format
There is no "correct" format for documents, but there is a correct format for documents you intend to feed to a language model. Markdown is the de facto standard: every major LLM was trained on it, every API accepts it, and every model produces better output from it. PDF is for printing. Markdown is for AI.
Why Markdown is the lingua franca of LLMs
Across OpenAI, Anthropic, Google DeepMind, Meta, Mistral, and the open-source long tail, training corpora are dominated by Markdown-flavoured text: README files, documentation sites, blog posts, GitHub wikis. The result is that every modern model recognises the same handful of cues: # means heading, - means list item, fenced code blocks are inviolable, tables are tables.
None of those cues exist in a PDF. PDF is a sequence of glyphs at coordinates; the structure has to be inferred. Inference costs tokens (the model thinks about layout instead of content) and introduces errors (the model gets the layout wrong). Markdown skips both costs.
Model-specific guides
The savings and best practices vary by model. We maintain a guide per major destination:
ChatGPT: token economics on GPT-4o / GPT-5 / o-series
Claude: Sonnet 4.6 and Opus 4.7 with the 200k context window
Gemini: 1M context on 2.5 Pro and AI Studio workflow
Markdown, by a wide margin. It carries semantic structure (headings, lists, code, tables) in a form every major LLM was trained to recognise, while consuming far fewer tokens than HTML, RTF, or extracted-PDF text.
Why not just use plain text instead of Markdown?
Plain text loses every structural cue: the model can't tell a heading from a sentence or a code block from prose. That forces it to either guess (errors) or treat everything uniformly (worse retrieval). Markdown adds the cues back at near-zero token cost.
Markdown vs JSON: which is better for LLM context?
Markdown for human-readable documents, JSON for structured data. JSON is more verbose (every field name is repeated) and harder for the model to skim. Use JSON when you need precise field access and Markdown when you need narrative context.
Do LLMs actually understand Markdown formatting?
Yes, fluently. Modern LLMs have been trained on so much Markdown that they treat **bold**, ## headings, fenced code blocks, and pipe-tables as semantic features, not just text. They will also generate Markdown output by default if you ask.
What's the average token reduction from PDF to Markdown?
On clean digital PDFs, 30–60% reduction. On layout-heavy PDFs (multi-column papers, reports), 60–80%. On scanned PDFs that the model would otherwise OCR internally, often >95%. The win compounds when you remove repeating headers, footers, and page numbers.