What Windsurf does with a Markdown file, and where a PDF fits
Windsurf takes Markdown through an at-mention of the file, or a workspace folder the assistant has been allowed to read, and that is the door this page is about. The question worth asking is what a PDF looks like by the time it arrives, because Windsurf never sees the PDF itself. It sees whatever the conversion handed over, and nothing else.
Windsurf reads the heading tree to decide which part of a long file to load. What comes out of a PDF here is headings where the document had headings, lists where it had lists, and pipe tables where it had ruled tables, so the structure Windsurf is looking for is actually present instead of merely implied. That is the entire trick, and there is nothing clever underneath it.
What survives from a PDF into Windsurf
The conversion drops running headers, page numbers, watermarks and the seam between two columns before Windsurf ever sees them, because none of it carries meaning into Windsurf and all of it costs something. What it keeps is the part you would otherwise have retyped by hand.
- Both heading levels come through, so the document keeps its own outline, and Windsurf receives it as structure rather than as something to infer.
- Ruled tables are rebuilt as aligned Markdown tables rather than flattened into prose, which is the part Windsurf would otherwise have to reconstruct on its own.
- The reading order survives a two column spread instead of interleaving the columns, so nothing downstream of Windsurf has to guess at it.
What Windsurf never gets back from a PDF
Every conversion costs something, and a page listing only the gains is a sales page. Here is what a PDF gives up on the way to Windsurf, stated plainly enough that you can decide against it.
- Page breaks stop existing, and any cross reference that said "see page 14" now points at nothing, and Windsurf will not get it back from anywhere.
- Fonts, margins and exact spacing are gone, because none of them carried meaning, so nothing you build in Windsurf should depend on it.
- Text that was only ever a picture of text is read by a model, not copied, which is a real cost inside Windsurf and not a rounding error.
The PDF to Markdown step behind this Windsurf page was measured
Gemini 2.5 Flash Lite does the conversion, and it was proven end to end on 30 August 2026: a PDF built for the occasion and never converted before came back with both heading levels, the witness word, and the ruled table rebuilt in aligned Markdown. That is one run on one day and not an average, which is exactly why the date travels with the claim instead of being left out of it. The same output is what Windsurf would have received.
A scanned PDF is read by a model rather than copied, and a model that misreads a digit does it with complete confidence. Windsurf has no way of telling you any of that from the inside, so it gets said here instead, on the page that sent you there.
Once the PDF is Markdown, what to do inside Windsurf
With the Markdown in hand, split a long conversion into several files named after their sections. That is a Windsurf habit rather than a conversion setting, and it is where most of the value of converting a PDF in the first place actually lands.
The constraint worth knowing before you start: the assistant loads part of a file, and which part it loads is not something you choose. It bites the same way whether the Markdown came out of a PDF or out of something else entirely, but it bites hardest on long documents, which is what a PDF usually is.
What we will not pretend about PDF in Windsurf
- Windsurf will not tell you which portion of the file it actually read, and converting a PDF first does not change that in the slightest.
- one very long file gets loaded partially and answered from confidently, which is a Windsurf behaviour rather than anything the PDF conversion did.
- The conversion costs half a credit and the price is shown before the PDF is sent, whether the Markdown ends up in Windsurf or nowhere at all.