Strip a PDF down to its text and nothing else. Useful for ctrl-F across documents, plugging into a search index, or feeding into a script. We strip layout furniture (page numbers, headers, footers) along the way so the text you get is content, not noise.
How "PDF to Text" actually works here
Every PDF goes through the same pipeline as our Markdown converter, but with the structural cues stripped out at the end. Headings stay (as plain lines), but lose their # markers; lists become indented text; tables flatten into tab-separated rows. The result is plain text — copy-paste-ready, search-indexable, ready for any tool that wants UTF-8 string input.
Scanned PDFs included
Scanned (image-only) PDFs are OCR'd automatically — you don't need a separate tool. Output looks the same: plain text, no layout markers. Quality of the OCR depends on the source: cleanly-scanned typed text comes out at 95%+ accuracy; phone-photographed pages depend on lighting.
Frequently asked questions
What's the difference between PDF to Text and PDF to Markdown?
PDF to Text gives you plain UTF-8 strings — flat, no structure. PDF to Markdown preserves headings, lists, tables, and code blocks. Use Text for search indexing or plain-text scripts; use Markdown for anything where structure matters (AI input, wikis, docs sites).
Does this work on scanned PDFs?
Yes — scanned PDFs (no text layer) are OCR'd automatically. The output is still plain text. Quality depends on source resolution; 200+ DPI scans typically come out at 95%+ accuracy.
Is the extracted text searchable?
It's plain UTF-8 — searchable in any tool that takes plain text (grep, full-text search engines, ctrl-F in any viewer). The way to make a PDF itself searchable is to convert and re-embed the text layer; that's a different workflow.
Are headers, footers, and page numbers removed?
Yes — repeating page-level furniture is detected and stripped. The extracted text is content only. If you want to keep them (rare), use the API with the strip_furniture=false flag.
Free PDF to Text vs paid OCR tools — what's the catch?
No catch — our free tier covers most personal use. The catch with "free OCR tools" elsewhere is usually watermarks, ads, file-size caps, or quietly uploading your file to a marketing pipeline. We don't do any of those.