What LLMs does with a Markdown file, and where an image fits
LLMs takes Markdown through whatever prompt box or file field the model happens to sit behind, and that is the door this page is about. The question worth asking is what an image looks like by the time it arrives, because LLMs never sees the image itself. It sees whatever the conversion handed over, and nothing else.
LLMs recognises the handful of cues every training corpus is full of, headings, list markers and fenced blocks. What comes out of an image here is the text of the image as Markdown, with a table rebuilt when the image held one, so the structure LLMs is looking for is actually present instead of merely implied. That is the entire trick, and there is nothing clever underneath it.
What survives from an image into LLMs
The conversion drops the background, the stamps, the shadows and the skew of the scan before LLMs ever sees them, because none of it carries meaning into LLMs and all of it costs something. What it keeps is the part you would otherwise have retyped by hand.
- The text comes out in reading order rather than in scan order, and LLMs receives it as structure rather than as something to infer.
- A visible table is rebuilt as a Markdown table instead of collapsing into a line of words, which is the part LLMs would otherwise have to reconstruct on its own.
- References and figures are read character by character, which is the part people retype by hand, so nothing downstream of LLMs has to guess at it.
What LLMs never gets back from an image
Every conversion costs something, and a page listing only the gains is a sales page. Here is what an image gives up on the way to LLMs, stated plainly enough that you can decide against it.
- Layout and colour are gone, and a colour that meant something meant it only to a human, and LLMs will not get it back from anywhere.
- Anything too small or too blurred to read is not returned, and not flagged either, so nothing you build in LLMs should depend on it.
- Certainty is gone: a model is reading, and a model can be wrong while sounding sure, which is a real cost inside LLMs and not a rounding error.
The Image to Markdown step behind this LLMs page was measured
Gemini 2.5 Flash Lite does the conversion, and it was proven end to end on 30 August 2026: a purchase order image built for the occasion, carrying a unique reference, came back with the reference, the quantity and the unit price all read correctly. That is one run on one day and not an average, which is exactly why the date travels with the claim instead of being left out of it. The same output is what LLMs would have received.
This capability returned nothing at all until 30 August 2026, and two probes failed that same day before the quota was restored. LLMs has no way of telling you any of that from the inside, so it gets said here instead, on the page that sent you there.
Once the image is Markdown, what to do inside LLMs
With the Markdown in hand, keep the Markdown as the source of truth and re-prompt from it rather than from the previous answer. That is a LLMs habit rather than a conversion setting, and it is where most of the value of converting an image in the first place actually lands.
The constraint worth knowing before you start: no model announces its token bill before it has already spent it. It bites the same way whether the Markdown came out of an image or out of something else entirely, but it bites hardest on long documents, which is what an image usually is.
What we will not pretend about Image in LLMs
- LLMs will not admit to having inferred a structure that was never there, and converting an image first does not change that in the slightest.
- structure a model has to guess is structure a model gets wrong on the long tail, which is a LLMs behaviour rather than anything the image conversion did.
- The conversion costs half a credit and the price is shown before the image is sent, whether the Markdown ends up in LLMs or nowhere at all.