Word to Markdown for LangChain: Better Than Docx2txt
Docx2txtLoader is the standard LangChain loader for Word documents, and it flattens every heading into plain text. Section structure disappears, downstream chunking has nothing to anchor on, and your retrieval surfaces context-less fragments. Pre-convert to Markdown and TextLoader plus MarkdownHeaderTextSplitter handles structure cleanly: each section becomes its own document with the heading path already in metadata.
The cleaner LangChain pipeline
The standard advice is Docx2txtLoader, which extracts text and discards structure. Fine for simple summarisation; painful for any pipeline that needs heading-aware chunking. The alternative: pre-convert each Word document on Word to Markdown, persist the .md, and use TextLoader from then on. Your loader becomes deterministic, your output is human-inspectable, and your splitter can rely on real section boundaries.
Pair with MarkdownHeaderTextSplitter
The single biggest win is the splitter. MarkdownHeaderTextSplitter chunks on actual heading boundaries instead of guessing: chunks correspond to document sections, the heading path lives in metadata, and retrieval-augmented prompts get free structural context. Combine with PDFs (PDF for LangChain), URLs (URL for LangChain), audio (Audio for LangChain), and video (Video for LangChain) for a multi-source pipeline.
Code example
# Local pipeline using LangChain on the .md you downloaded from mdisbetter.com.
# Install: pip install langchain-community langchain-text-splitters
from langchain_community.document_loaders import TextLoader
from langchain_text_splitters import MarkdownHeaderTextSplitter, RecursiveCharacterTextSplitter
# 1. Load the converted Word document
docs = TextLoader("vendor-contract-template.md", encoding="utf-8").load()
md_text = docs[0].page_content
# 2. Split on section headings — each ## becomes a chunk with section metadata
md_splitter = MarkdownHeaderTextSplitter(headers_to_split_on=[
("#", "document"),
("##", "section"), # ## 4. Termination
("###", "subsection"),
])
section_chunks = md_splitter.split_text(md_text)
# 3. Sub-split any over-budget section for your embedding model
splitter = RecursiveCharacterTextSplitter(chunk_size=1000, chunk_overlap=120)
chunks = splitter.split_documents(section_chunks)
# Each chunk's metadata['section'] is e.g. "4. Termination"
Frequently asked questions
MarkdownHeaderTextSplitter vs Docx2txtLoader for Word documents?
Docx2txtLoader returns flat text: heading structure flattened, chunking has to guess at boundaries. Pre-converting to Markdown and using MarkdownHeaderTextSplitter gives you section-aware chunks with heading metadata. Use the converter for the structure, then any standard LangChain text-handling primitive works downstream.
Can I attach section metadata to every chunk automatically?
Yes, automatic with MarkdownHeaderTextSplitter. The heading text (e.g. 4. Termination or 3.2 Force Majeure) becomes a metadata field on each chunk. Retrieval can filter by section; synthesis can cite by section heading.
How is this different from UnstructuredWordDocumentLoader?
UnstructuredWordDocumentLoader uses the unstructured library to extract richer metadata: it's a step up from Docx2txtLoader. The pre-conversion approach is even cleaner: you get Markdown you can inspect and hand-correct before ingestion, and the same TextLoader code works across all your converted source modalities (PDF, URL, audio, video, Word).
What about Word documents with embedded tables and lists?
The converter renders Word tables as Markdown tables (pipe syntax) and Word lists as Markdown lists. MarkdownHeaderTextSplitter respects table and list boundaries: chunks won't slice through a table mid-row. For documents heavy in structured data (compliance matrices, comparison tables), this is significantly cleaner than Docx2txtLoader's output.
Should I run conversion on a server or use the web tool?
For ad-hoc and progressive ingestion (a few documents at a time), the web tool at mdisbetter.com is the right surface: upload, download, drop into your pipeline. For automated batch ingestion of thousands of documents, run Pandoc locally in a script (pandoc input.docx -o output.md). Same Markdown output, different operational surfaces.