What Docusaurus does with a Markdown file, and where a legacy Word document fits
Docusaurus takes Markdown through a file under the docs directory, with an identifier and a sidebar position in its front matter, and that is the door this page is about. The question worth asking is what a legacy Word document looks like by the time it arrives, because Docusaurus never sees the document itself. It sees whatever the conversion handed over, and nothing else.
Docusaurus treats the file as MDX, which means characters that are harmless in plain Markdown are not harmless here. What comes out of a legacy Word document here is headings, bold, italics, nested lists and tables, in the order the document had them, so the structure Docusaurus is looking for is actually present instead of merely implied. That is the entire trick, and there is nothing clever underneath it.
What survives from a legacy Word document into Docusaurus
The conversion drops field codes, revision marks, print settings and the styles nobody ever applied before Docusaurus ever sees them, because none of it carries meaning into Docusaurus and all of it costs something. What it keeps is the part you would otherwise have retyped by hand.
- Heading levels and list nesting survive, which is what makes the document navigable afterwards, and Docusaurus receives it as structure rather than as something to infer.
- Bold and italics survive, so emphasis that carried meaning still carries it, which is the part Docusaurus would otherwise have to reconstruct on its own.
- Tables survive as tables rather than as runs of tab separated text, so nothing downstream of Docusaurus has to guess at it.
What Docusaurus never gets back from a legacy Word document
Every conversion costs something, and a page listing only the gains is a sales page. Here is what a legacy Word document gives up on the way to Docusaurus, stated plainly enough that you can decide against it.
- Tracked changes and comments do not come through, and they are often the interesting part, and Docusaurus will not get it back from anywhere.
- Page setup and fonts are gone, along with anything that depended on where a page ended, so nothing you build in Docusaurus should depend on it.
- Objects embedded rather than written, such as a spreadsheet dropped into the page, are not unpacked, which is a real cost inside Docusaurus and not a rounding error.
The Word to Markdown step behind this Docusaurus page was measured
LibreOffice 24.2 does the conversion, and it was proven end to end on 13 August 2026: binary .doc, .rtf and .odt all came back with headings, bold, italics, lists and tables intact, and a 2 MB RTF of 4000 sections and 80 tables came back in 3.1 seconds. That is one run on one day and not an average, which is exactly why the date travels with the claim instead of being left out of it. The same output is what Docusaurus would have received.
Above 25 MB the answer is a refusal, and a document LibreOffice cannot open is refused outright rather than returned empty. Docusaurus has no way of telling you any of that from the inside, so it gets said here instead, on the page that sent you there.
Once the document is Markdown, what to do inside Docusaurus
With the Markdown in hand, set the identifier and sidebar position before the first build, because changing an identifier later breaks every link to it. That is a Docusaurus habit rather than a conversion setting, and it is where most of the value of converting a legacy Word document in the first place actually lands.
The constraint worth knowing before you start: an unescaped angle bracket or curly brace stops the build with an error pointing at the wrong line. It bites the same way whether the Markdown came out of a legacy Word document or out of something else entirely, but it bites hardest on long documents, which is what a legacy Word document usually is.
What we will not pretend about Word in Docusaurus
- Docusaurus will not escape the characters MDX chokes on for you, and converting a legacy Word document first does not change that in the slightest.
- a converted document full of angle brackets is exactly the input MDX handles worst, which is a Docusaurus behaviour rather than anything the document conversion did.
- The conversion costs half a credit and the price is shown before the document is sent, whether the Markdown ends up in Docusaurus or nowhere at all.