PDF to Markdown for LlamaIndex: Structured Ingestion
LlamaIndex's SimpleDirectoryReader will happily ingest a PDF, but its default node-parsing on PDF text is generic: sentence splitting on whatever the PDF parser hands back. Pre-convert to Markdown and you can use MarkdownNodeParser, which builds a node tree that mirrors the document's real heading hierarchy.
Why MarkdownNodeParser changes the math
The default LlamaIndex flow on PDFs flattens everything to text and then sentence-splits. That's fine for short documents but disastrous for anything structured: a 60-page report becomes a long list of sentences with no hint that section 4.2 is part of chapter 4. Retrieval then surfaces orphan sentences instead of coherent passages.
MarkdownNodeParser builds a hierarchical node graph: H1 nodes contain H2 nodes contain H3 nodes contain paragraph nodes. Retrieval can climb the hierarchy, your prompts can include parent context, and your re-ranker has structural signal to work with.
Code example
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex
from llama_index.core.node_parser import MarkdownNodeParser
# 1. Load pre-converted Markdown files
documents = SimpleDirectoryReader(
input_dir="./markdown_docs",
required_exts=[".md"],
).load_data()
# 2. Parse into hierarchical nodes via the Markdown structure
parser = MarkdownNodeParser()
nodes = parser.get_nodes_from_documents(documents)
# 3. Index — each node now carries its heading path as metadata
index = VectorStoreIndex(nodes)
query_engine = index.as_query_engine(similarity_top_k=5)
Frequently asked questions
MarkdownNodeParser vs SentenceSplitter for PDF content?
MarkdownNodeParser respects document hierarchy and gives nodes structural metadata. SentenceSplitter just chops on sentence boundaries: fine for prose, terrible for technical documents where section context matters. For converted PDF content, always prefer MarkdownNodeParser.
How do I set up a LlamaIndex ingestion pipeline with Markdown?
Standard pattern: load .md files with SimpleDirectoryReader (required_exts=[".md"]), parse with MarkdownNodeParser, build a VectorStoreIndex over the nodes. The whole pipeline is ~10 lines and produces noticeably better retrieval than the equivalent on raw PDF.
Does Markdown preserve metadata for LlamaIndex nodes?
Yes: MarkdownNodeParser automatically promotes the heading path to node metadata, so each node knows it lives under "Chapter 4 > Section 4.2 > Methodology". Your queries and re-rankers can use that metadata directly.
Can I build hierarchical nodes from Markdown headings?
That's exactly what MarkdownNodeParser does. The resulting nodes form a tree that mirrors your document's table of contents, which means you can do parent-child retrieval, summary-based routing, or auto-merging-retriever patterns out of the box.
SimpleDirectoryReader vs pre-converted Markdown: which is better?
For one-off ingestion, SimpleDirectoryReader on PDFs is simpler. For anything you'll re-index more than once, pre-convert to Markdown: your indexing time drops, your chunks become deterministic, and your output is human-inspectable for debugging.