DOCX is Word in the OOXML format, with headings, tables and named styles. TXT is plain text, with no markup of any kind. pandoc sits between the two, and this page says exactly what happens to a DOCX file on the way to becoming a TXT one.
What runs when a DOCX file becomes TXT
pandoc in two passes, on a machine we rent and watch. Your DOCX file is uploaded once, pandoc runs once, the TXT comes back, and neither file is kept. DOCX to TXT is one of the 3649 pairs that engine was probed on with a real file, which is why it has a page here and why the pairs the probe could not prove do not.
What survives from the DOCX into the TXT
Lists from the DOCX file, nesting included, plus code blocks, block quotes and links, all re-expressed in TXT.
Images: pandoc extracts them out of the DOCX file into a working directory and re-embeds them in the TXT, so nothing ends up pointing at a file that no longer exists.
What DOCX to TXT costs you
Page layout. DOCX carries some and TXT carries none, so margins, page breaks, headers and footers have nowhere to land and are simply not written.
Tables. The TXT writer cannot carry one, and the probe caught this the hard way: a CSV file, which is nothing but a table, came back as a well formed and completely empty TXT with a zero return code. The engine now refuses that case outright rather than handing you the emptiness.
Where DOCX and TXT files come from
DOCX. Word writes it, and so does Google Docs on export. It is ISO/IEC 29500, a zip of XML, which is what lets a program read the document without opening Word.
If a model is going to read the TXT
Structure is what survives from the DOCX file, and structure is what a model needs. DOCX in, TXT out, with the heading tree intact instead of flattened into one long paragraph.
DOCX to TXT, measured rather than promised
Running DOCX to TXT against the real engine with a real file, pandoc wrote 264 bytes of TXT in 0.1 seconds, machine otherwise idle. That is one DOCX file on one day and not an average, which is why the number is given together with the file that produced it.
The DOCX to TXT verdict was reached by reading the text back, meaning the output decodes as readable text and still carries the witness word, which is the only check a plain text target allows. Which check was used matters, because they do not all prove the same thing, and a status code of 200 proves nothing whatsoever about whether the TXT file has anything inside it.
The same DOCX file, sent somewhere else
Other outputs the probe measured out of a DOCX file, so the cost of choosing TXT can be read against something. Sizes do not compare across engines, because each family was probed with its own DOCX witness file.
DOCX to PDF: 11,917 bytes in 1.2 seconds, verified by a format signature.
DOCX to MD: 266 bytes in 0.1 seconds, verified by a pandoc round trip.
DOCX to EPUB: 5,040 bytes in 0.1 seconds, verified by a pandoc round trip.
DOCX to HTML: 4,018 bytes in 0.1 seconds, verified by a pandoc round trip.
DOCX to RTF: 4,205 bytes in 1.1 seconds, verified by reading the internal type and finding a witness word.
DOCX to ODT: 12,806 bytes in 1.2 seconds, verified by reading the internal type and finding a witness word.
Other ways into TXT, and what they measured
Among the published routes into TXT, DOCX is the third largest output of the 11 measured. The witness files differ, so this ranks the probe run and not your document.
PDF to TXT: 219 bytes in 0.8 seconds.
DOC to TXT: 268 bytes in 1.1 seconds.
AZW3 to TXT: 219 bytes in 0.6 seconds.
CSV to TXT: 114 bytes in under a tenth of a second.
HTML to TXT: 270 bytes in under a tenth of a second.
What we will not pretend about DOCX to TXT
A DOCX file over 25 MB is refused before the upload finishes rather than after it, so you do not wait for a rejection.
A DOCX to TXT run that passes 60 seconds is killed, and the pandoc process is killed with it. A run left behind would sit on one of the machine's two cores until somebody noticed.
32 of the 935 pairs probed in the family that serves DOCX to TXT failed, and this pair is not one of them. They fail for reasons worth knowing: a writer that cannot carry what the document is made of, or output the engine could not read back. They are counted here rather than hidden, because a pair that fails quietly is worse than one that fails loudly.
pandoc never writes the TXT through a TeX engine here, and it never writes PDF at all: the eleven engines it would need are not on this machine and a TeX distribution weighs several gigabytes. Anyone who wants a PDF out of a DOCX file goes through LibreOffice or calibre, both of which genuinely can.
One DOCX file at a time, chosen in the browser. There is nothing else to set up and nothing else on offer.
DOCX to TXT: what people ask
What actually converts my DOCX file to TXT?
pandoc does it, in two passes, on a machine we rent and watch. Not a browser trick and not somebody else service: the DOCX file is uploaded once, pandoc runs once, the TXT comes back, and neither file is kept afterwards.
How long does DOCX to TXT take?
On the file the probe used, pandoc took 0.1 seconds and wrote 264 bytes of TXT. That is one real measurement on one real DOCX file, not an average and not a promise about yours: a larger DOCX takes longer, and past 60 seconds the run is stopped.
What do I lose going from DOCX to TXT?
The one to know about first: Page layout. DOCX carries some and TXT carries none, so margins, page breaks, headers and footers have nowhere to land and are simply not written.
Are the tables in my DOCX file still tables in the TXT?
No, and that is a property of TXT rather than a defect of this path. TXT cannot hold a table, so a DOCX file built around one is the wrong candidate for this pair.
Is DOCX to TXT free?
There is a free allowance every month, and one DOCX to TXT conversion costs half a credit against it. When the allowance runs out the tool says so and stops, rather than quietly handing you a worse TXT.