What LlamaIndex does with a Markdown file, and where an image fits
LlamaIndex takes Markdown through the Markdown node parser, which turns a file into nodes before anything is embedded, and that is the door this page is about. The question worth asking is what an image looks like by the time it arrives, because LlamaIndex never sees the image itself. It sees whatever the conversion handed over, and nothing else.
LlamaIndex carries the heading path down into node metadata, so a retrieved node still knows where it used to live. What comes out of an image here is the text of the image as Markdown, with a table rebuilt when the image held one, so the structure LlamaIndex is looking for is actually present instead of merely implied. That is the entire trick, and there is nothing clever underneath it.
What survives from an image into LlamaIndex
The conversion drops the background, the stamps, the shadows and the skew of the scan before LlamaIndex ever sees them, because none of it carries meaning into LlamaIndex and all of it costs something. What it keeps is the part you would otherwise have retyped by hand.
- The text comes out in reading order rather than in scan order, and LlamaIndex receives it as structure rather than as something to infer.
- A visible table is rebuilt as a Markdown table instead of collapsing into a line of words, which is the part LlamaIndex would otherwise have to reconstruct on its own.
- References and figures are read character by character, which is the part people retype by hand, so nothing downstream of LlamaIndex has to guess at it.
What LlamaIndex never gets back from an image
Every conversion costs something, and a page listing only the gains is a sales page. Here is what an image gives up on the way to LlamaIndex, stated plainly enough that you can decide against it.
- Layout and colour are gone, and a colour that meant something meant it only to a human, and LlamaIndex will not get it back from anywhere.
- Anything too small or too blurred to read is not returned, and not flagged either, so nothing you build in LlamaIndex should depend on it.
- Certainty is gone: a model is reading, and a model can be wrong while sounding sure, which is a real cost inside LlamaIndex and not a rounding error.
The Image to Markdown step behind this LlamaIndex page was measured
Gemini 2.5 Flash Lite does the conversion, and it was proven end to end on 30 August 2026: a purchase order image built for the occasion, carrying a unique reference, came back with the reference, the quantity and the unit price all read correctly. That is one run on one day and not an average, which is exactly why the date travels with the claim instead of being left out of it. The same output is what LlamaIndex would have received.
This capability returned nothing at all until 30 August 2026, and two probes failed that same day before the quota was restored. LlamaIndex has no way of telling you any of that from the inside, so it gets said here instead, on the page that sent you there.
Once the image is Markdown, what to do inside LlamaIndex
With the Markdown in hand, read the node text before the first embedding call, because that is the only cheap moment to catch a bad split. That is a LlamaIndex habit rather than a conversion setting, and it is where most of the value of converting an image in the first place actually lands.
The constraint worth knowing before you start: nodes inherit whatever noise the loader left behind, and nothing downstream removes it. It bites the same way whether the Markdown came out of an image or out of something else entirely, but it bites hardest on long documents, which is what an image usually is.
What we will not pretend about Image in LlamaIndex
- LlamaIndex will not tell you that two nodes are near duplicates of each other, and converting an image first does not change that in the slightest.
- a parser fed reflowed text produces nodes that break in the middle of a clause, which is a LlamaIndex behaviour rather than anything the image conversion did.
- The conversion costs half a credit and the price is shown before the image is sent, whether the Markdown ends up in LlamaIndex or nowhere at all.