HTML is the markup of the web, with native headings and tables. DOCX is Word in the OOXML format, with headings, tables and named styles. pandoc sits between the two, and this page says exactly what happens to an HTML file on the way to becoming a DOCX one.
What runs when an HTML file becomes DOCX
pandoc in two passes, on a machine we rent and watch. Your HTML file is uploaded once, pandoc runs once, the DOCX comes back, and neither file is kept. HTML to DOCX is one of the 3649 pairs that engine was probed on with a real file, which is why it has a page here and why the pairs the probe could not prove do not.
What survives from the HTML into the DOCX
The heading hierarchy. pandoc parses the HTML file into its own document tree and writes that tree out as DOCX, so a level two heading is still a level two heading and not merely bold text.
Tables, rebuilt in DOCX syntax rather than screenshotted. That is the difference between a table a machine can read and a picture of one, and it is the whole reason to send an HTML file through pandoc rather than through a printer.
Lists from the HTML file, nesting included, plus code blocks, block quotes and links, all re-expressed in DOCX.
Images: pandoc extracts them out of the HTML file into a working directory and re-embeds them in the DOCX, so nothing ends up pointing at a file that no longer exists.
What HTML to DOCX costs you
Nothing that was ever there, which is the honest answer for this pair: an HTML file has no page layout to lose. The DOCX you get uses the default Word template, so it is a plain document rather than a designed one.
Where HTML and DOCX files come from
DOCX. Word writes it, and so does Google Docs on export. It is ISO/IEC 29500, a zip of XML, which is what lets a program read the document without opening Word.
If a model is going to read the DOCX
DOCX is a good destination for a person and a mediocre one for a model: the structure is real but buried in XML with a great deal of packaging around it. If the next reader of this HTML file is an assistant rather than a colleague, HTML to Markdown is the shorter path and costs fewer tokens for exactly the same content.
HTML to DOCX, measured rather than promised
Running HTML to DOCX against the real engine with a real file, pandoc wrote 10,277 bytes of DOCX in 0.1 seconds, machine otherwise idle. That is one HTML file on one day and not an average, which is why the number is given together with the file that produced it.
The HTML to DOCX verdict was reached by a pandoc round trip, meaning the output was read back by pandoc and had to still contain a witness word planted in the source, which proves the content travelled and not merely the container. Which check was used matters, because they do not all prove the same thing, and a status code of 200 proves nothing whatsoever about whether the DOCX file has anything inside it.
The same HTML file, sent somewhere else
Other outputs the probe measured out of an HTML file, so the cost of choosing DOCX can be read against something. Sizes do not compare across engines, because each family was probed with its own HTML witness file.
HTML to PDF: 25,178 bytes in 1.9 seconds, verified by a format signature.
HTML to MD: 371 bytes in under a tenth of a second, verified by a pandoc round trip.
HTML to EPUB: 5,027 bytes in 0.1 seconds, verified by a pandoc round trip.
HTML to TXT: 270 bytes in under a tenth of a second, verified by reading the text back.
HTML to RTF: 775 bytes in under a tenth of a second, verified by a pandoc round trip.
HTML to ODT: 7,629 bytes in 0.1 seconds, verified by a pandoc round trip.
Other ways into DOCX, and what they measured
Among the published routes into DOCX, HTML is the fifth largest output of the 12 measured. The witness files differ, so this ranks the probe run and not your document.
TXT to DOCX: 9,900 bytes in 0.1 seconds.
DOC to DOCX: 5,462 bytes in 1.1 seconds.
TEX to DOCX: 10,172 bytes in 0.1 seconds.
AZW3 to DOCX: 23,936 bytes in 0.6 seconds.
CSV to DOCX: 9,945 bytes in 0.1 seconds.
What we will not pretend about HTML to DOCX
An HTML file over 25 MB is refused before the upload finishes rather than after it, so you do not wait for a rejection.
An HTML to DOCX run that passes 60 seconds is killed, and the pandoc process is killed with it. A run left behind would sit on one of the machine's two cores until somebody noticed.
32 of the 935 pairs probed in the family that serves HTML to DOCX failed, and this pair is not one of them. They fail for reasons worth knowing: a writer that cannot carry what the document is made of, or output the engine could not read back. They are counted here rather than hidden, because a pair that fails quietly is worse than one that fails loudly.
pandoc never writes the DOCX through a TeX engine here, and it never writes PDF at all: the eleven engines it would need are not on this machine and a TeX distribution weighs several gigabytes. Anyone who wants a PDF out of an HTML file goes through LibreOffice or calibre, both of which genuinely can.
One HTML file at a time, chosen in the browser. There is nothing else to set up and nothing else on offer.
HTML to DOCX: what people ask
What actually converts my HTML file to DOCX?
pandoc does it, in two passes, on a machine we rent and watch. Not a browser trick and not somebody else service: the HTML file is uploaded once, pandoc runs once, the DOCX comes back, and neither file is kept afterwards.
How long does HTML to DOCX take?
On the file the probe used, pandoc took 0.1 seconds and wrote 10,277 bytes of DOCX. That is one real measurement on one real HTML file, not an average and not a promise about yours: a larger HTML takes longer, and past 60 seconds the run is stopped.
What do I lose going from HTML to DOCX?
The one to know about first: Nothing that was ever there, which is the honest answer for this pair: an HTML file has no page layout to lose. The DOCX you get uses the default Word template, so it is a plain document rather than a designed one.
Are the tables in my HTML file still tables in the DOCX?
Yes. pandoc reads the HTML tables into its own document tree and writes them back out in DOCX syntax, so what you get is a table a machine can parse rather than a picture of one.
Is HTML to DOCX free?
There is a free allowance every month, and one HTML to DOCX conversion costs half a credit against it. When the allowance runs out the tool says so and stops, rather than quietly handing you a worse DOCX.