Clean PDF for LLM Context: Remove Noise, Keep Structure
Most "clean my PDF for AI" tools strip too much (losing structure) or not enough (keeping page numbers and footers). The right cleaning is opinionated: drop everything an LLM doesn't need, keep everything it does. Markdown conversion does this by design.
How much a typical PDF wastes
We benchmarked 10 representative documents: academic papers, product manuals, financial reports, legal contracts, slide decks. Average token reduction from PDF text to Markdown: 68%. Worst case: 41% (a clean digital paper with minimal furniture). Best case: 96% (a scanned, multi-column report where the OCR text itself was mostly noise).
Where does the saving come from? Roughly 25% from removing repeating headers, footers and page numbers; 30% from collapsing whitespace and normalising encoding; the rest from dropping invisible glyphs, watermarks, and broken column boundaries that produced duplicated content in extraction.
What to keep, what to strip
Keep: headings, lists, code, tables, links, math notation, and paragraph breaks that respect the document's argument structure.
Strip: page numbers, repeating headers and footers, watermarks, "Page X of Y" markers, copyright lines on every page, and decorative separators. None of them help the LLM; all of them cost tokens.
Frequently asked questions
How many tokens does a typical PDF waste on formatting noise?
On clean digital PDFs, 30–50% waste. On layout-heavy reports and multi-column papers, 60–80%. On scanned PDFs that need internal OCR, often >90%. Averaged across our 10-document benchmark: 68%.
What elements in a PDF are "noise" for LLMs?
Page numbers, repeating headers and footers, watermarks, "Page X of Y" markers, copyright lines per page, sidebar callouts whose content doesn't belong to the main flow, and column-break artefacts. None of them carry meaning the LLM needs.
Does cleaning a PDF remove important information?
Done well, no: only layout furniture is removed, while content (headings, paragraphs, lists, tables, code, math) is preserved. Done crudely (regex strip everything that looks like a number), yes: you can lose page references, equation numbers, or section IDs. Markdown conversion does the well-done version.
Cleaning vs full conversion: what's the difference?
"Cleaning" leaves the file as PDF and just removes furniture: useful if a downstream tool requires PDF input. "Full conversion" produces Markdown the LLM can read directly. For LLM context, full conversion is always better; cleaning alone still leaves the model parsing layout.
How do headers, footers, and page numbers affect LLM output?
Repeating headers and footers create false self-similarity (every chunk looks slightly like every other), which confuses retrieval in RAG. Page numbers leak into citations ("the document mentions page 14"). Stripping all three before context injection consistently improves answer accuracy.