PDF is a frozen page layout with no heading structure a machine can rely on. DOCX is Word in the OOXML format, with headings, tables and named styles. calibre sits between the two, and this page says exactly what happens to a PDF file on the way to becoming a DOCX one.
What runs when a PDF file becomes DOCX
calibre the ebook-convert binary, on a machine we rent and watch. Your PDF file is uploaded once, calibre runs once, the DOCX comes back, and neither file is kept. PDF to DOCX is one of the 3649 pairs that engine was probed on with a real file, which is why it has a page here and why the pairs the probe could not prove do not.
What survives from the PDF into the DOCX
The reading order and the chapter split of the PDF file. calibre rebuilds it as a book on the way to DOCX, rather than as a pile of pages.
Cover art and book metadata, which calibre carries from the PDF container into the DOCX one.
What PDF to DOCX costs you
Fixed layout. calibre reflows the PDF text by design, because a DOCX is meant to be read at any font size on any screen width.
Anything the PDF did not spell out. calibre reads the text layer of the PDF file; a scanned page has no text layer, so a scanned PDF comes back as an empty DOCX, and no engine on this machine can warn you beforehand.
Where PDF and DOCX files come from
DOCX. Word writes it, and so does Google Docs on export. It is ISO/IEC 29500, a zip of XML, which is what lets a program read the document without opening Word.
If a model is going to read the DOCX
DOCX is a good destination for a person and a mediocre one for a model: the structure is real but buried in XML with a great deal of packaging around it. If the next reader of this PDF file is an assistant rather than a colleague, PDF to Markdown is the shorter path and costs fewer tokens for exactly the same content.
PDF to DOCX, measured rather than promised
Running PDF to DOCX against the real engine with a real file, calibre wrote 23,905 bytes of DOCX in 0.9 seconds, machine otherwise idle. That is one PDF file on one day and not an average, which is why the number is given together with the file that produced it.
The PDF to DOCX verdict was reached by a format signature, meaning at least four bytes unique to the format were found at the right offset. Which check was used matters, because they do not all prove the same thing, and a status code of 200 proves nothing whatsoever about whether the DOCX file has anything inside it.
The same PDF file, sent somewhere else
Other outputs the probe measured out of a PDF file, so the cost of choosing DOCX can be read against something. Sizes do not compare across engines, because each family was probed with its own PDF witness file.
PDF to EPUB: 26,909 bytes in 0.9 seconds, verified by a format signature.
PDF to JPG: 5,782 bytes in 0.1 seconds, verified by a format signature.
PDF to PNG: 1,294 bytes in 0.1 seconds, verified by a format signature.
PDF to WEBP: 906 bytes in 0.1 seconds, verified by a format signature.
PDF to TXT: 219 bytes in 0.8 seconds, verified by reading the text back.
PDF to GIF: 9,463 bytes in 0.1 seconds, verified by a format signature.
Other ways into DOCX, and what they measured
Among the published routes into DOCX, PDF is the third largest output of the 12 measured. The witness files differ, so this ranks the probe run and not your document.
EPUB to DOCX: 10,326 bytes in 0.1 seconds.
ODT to DOCX: 5,514 bytes in 1.1 seconds.
RTF to DOCX: 5,422 bytes in 1.1 seconds.
HTML to DOCX: 10,277 bytes in 0.1 seconds.
TXT to DOCX: 9,900 bytes in 0.1 seconds.
What we will not pretend about PDF to DOCX
A PDF file over 50 MB is refused before the upload finishes rather than after it, so you do not wait for a rejection.
A PDF to DOCX run that passes 120 seconds is killed, and the calibre process is killed with it. A run left behind would sit on one of the machine's two cores until somebody noticed.
The PDF file used to prove this pair was short. calibre is allowed 120 seconds and a long book can genuinely need all of them, so the DOCX timing below is a measurement of one small file and not a guarantee about your thousand page one.
One PDF file at a time, chosen in the browser. There is nothing else to set up and nothing else on offer.
PDF to DOCX: what people ask
What actually converts my PDF file to DOCX?
calibre does it, the ebook-convert binary, on a machine we rent and watch. Not a browser trick and not somebody else service: the PDF file is uploaded once, calibre runs once, the DOCX comes back, and neither file is kept afterwards.
How long does PDF to DOCX take?
On the file the probe used, calibre took 0.9 seconds and wrote 23,905 bytes of DOCX. That is one real measurement on one real PDF file, not an average and not a promise about yours: a larger PDF takes longer, and past 120 seconds the run is stopped.
What do I lose going from PDF to DOCX?
The one to know about first: Fixed layout. calibre reflows the PDF text by design, because a DOCX is meant to be read at any font size on any screen width.
Does the table of contents survive from PDF to DOCX?
It does, as long as your PDF file has real headings rather than text that merely looks like headings. calibre builds the DOCX contents from the heading structure it finds, and hand-formatted big bold lines are not a structure.
Is PDF to DOCX free?
There is a free allowance every month, and one PDF to DOCX conversion costs half a credit against it. When the allowance runs out the tool says so and stops, rather than quietly handing you a worse DOCX.