What Embeddings does with a Markdown file, and where an image fits
Embeddings takes Markdown through a chunk of text and nothing else, since that is the entire input, and that is the door this page is about. The question worth asking is what an image looks like by the time it arrives, because Embeddings never sees the image itself. It sees whatever the conversion handed over, and nothing else.
Embeddings averages meaning across whatever it is handed, furniture included. What comes out of an image here is the text of the image as Markdown, with a table rebuilt when the image held one, so the structure Embeddings is looking for is actually present instead of merely implied. That is the entire trick, and there is nothing clever underneath it.
What survives from an image into Embeddings
The conversion drops the background, the stamps, the shadows and the skew of the scan before Embeddings ever sees them, because none of it carries meaning into Embeddings and all of it costs something. What it keeps is the part you would otherwise have retyped by hand.
- The text comes out in reading order rather than in scan order, and Embeddings receives it as structure rather than as something to infer.
- A visible table is rebuilt as a Markdown table instead of collapsing into a line of words, which is the part Embeddings would otherwise have to reconstruct on its own.
- References and figures are read character by character, which is the part people retype by hand, so nothing downstream of Embeddings has to guess at it.
What Embeddings never gets back from an image
Every conversion costs something, and a page listing only the gains is a sales page. Here is what an image gives up on the way to Embeddings, stated plainly enough that you can decide against it.
- Layout and colour are gone, and a colour that meant something meant it only to a human, and Embeddings will not get it back from anywhere.
- Anything too small or too blurred to read is not returned, and not flagged either, so nothing you build in Embeddings should depend on it.
- Certainty is gone: a model is reading, and a model can be wrong while sounding sure, which is a real cost inside Embeddings and not a rounding error.
The Image to Markdown step behind this Embeddings page was measured
Gemini 2.5 Flash Lite does the conversion, and it was proven end to end on 30 August 2026: a purchase order image built for the occasion, carrying a unique reference, came back with the reference, the quantity and the unit price all read correctly. That is one run on one day and not an average, which is exactly why the date travels with the claim instead of being left out of it. The same output is what Embeddings would have received.
This capability returned nothing at all until 30 August 2026, and two probes failed that same day before the quota was restored. Embeddings has no way of telling you any of that from the inside, so it gets said here instead, on the page that sent you there.
Once the image is Markdown, what to do inside Embeddings
With the Markdown in hand, check a few nearest neighbours by hand before embedding an entire corpus. That is a Embeddings habit rather than a conversion setting, and it is where most of the value of converting an image in the first place actually lands.
The constraint worth knowing before you start: the vector has no idea which words were content and which were layout. It bites the same way whether the Markdown came out of an image or out of something else entirely, but it bites hardest on long documents, which is what an image usually is.
What we will not pretend about Image in Embeddings
- Embeddings will not let you inspect a vector, only compare it to another one, and converting an image first does not change that in the slightest.
- page furniture repeated across chunks pulls unrelated passages toward each other, which is a Embeddings behaviour rather than anything the image conversion did.
- The conversion costs half a credit and the price is shown before the image is sent, whether the Markdown ends up in Embeddings or nowhere at all.