PDF to Markdown for Vector Databases: Embedding-Ready Text
Whatever vector store you use (Pinecone, Chroma, Weaviate, Qdrant, pgvector), its retrieval quality is bounded by the embeddings you feed it, and your embeddings are bounded by your input text. Raw PDF chunks embed layout noise alongside content. Markdown chunks embed only content.
The chunk-quality bottleneck nobody talks about
Vector database benchmarks usually compare retrieval algorithms: flat index vs HNSW, dot vs cosine, ANN parameters. They almost never compare input quality, even though it's the single largest source of variance. Try the same query against the same vector DB on the same documents but with one corpus indexed from raw PDF text and the other from clean Markdown: top-K accuracy differs by 10–25 points on most public datasets.
What to keep in your embedding chunks
Markdown headings, prose, lists, code blocks, and tables. Strip page numbers, running headers, and footers before chunking: they make every chunk look slightly different from every other chunk in ways that have nothing to do with content. Keep heading paths as chunk metadata so you can filter retrieval by section without reembedding.
Frequently asked questions
How does Markdown input improve embedding quality?
Embeddings reflect everything in the input, including layout artefacts. Removing page numbers, headers, and column noise lets the embedding cluster on actual semantic content. The result is denser, more discriminative vectors and 10–25% better top-K retrieval on most corpora.
What chunk size works best for vector databases?
For OpenAI text-embedding-3-large, target 600–1000 tokens. For Cohere embed-v3, similar. Smaller chunks improve precision, larger chunks improve recall: pick based on whether your downstream task is "find the exact passage" (smaller) or "summarise relevant material" (larger).
Should I keep Markdown formatting in embeddings?
Mostly yes. Headings should be preserved (they're semantic anchors). Code blocks should stay fenced (so they embed as code, not prose). You can drop bold/italic/link syntax if you want: it makes a small positive difference on dense retrieval, no difference on lexical retrieval.
How do I index Markdown chunks in Pinecone?
Standard flow: chunk by headers, embed each chunk with your model of choice, upsert to Pinecone with the heading path and any other useful filters as metadata. Pinecone's metadata filtering then lets you scope retrieval to specific sections without re-embedding.
Batch indexing: how many Markdown chunks per request?
Most vector DBs accept 100–1000 vectors per upsert. Embedding APIs (OpenAI, Cohere, Voyage) typically batch 96–128 inputs per call. Match the smaller of the two for the simplest pipeline; parallelise if you're ingesting millions of chunks.