The Magic Codex

PDF Text Extractor

Drop in a PDF, read every page's text. Your file never leaves your browser — extraction happens entirely on your device.

Drop your PDF here

or tap to choose a file · stays on your device, nothing uploaded

No PDF loaded yet — drop one above to begin.

How to use the PDF Text Extractor

  1. Drop a PDF onto the dashed box (or tap it to browse your files). It never leaves your device.
  2. Flip through the pages with Prev/Next or jump straight to a page number — watch the word and character counts tally up.
  3. Copy all to grab every page at once, or Download .txt for a labeled plain-text file you can keep.

Why extraction beats retyping

PDFs are built for printing, not for copying — text inside them is often fragmented, out of order, or locked behind a viewer that fights selection. This tool walks the document's actual text layer page by page and reassembles it into clean, selectable text. Students use it to pull quotes for citations, researchers to feed papers into notes, and everyone else to rescue text from PDFs that won't let you highlight.

Scanned PDFs, explained

If the tool reports no embedded text, your PDF is almost certainly a scan: someone photographed or flatbed-scanned paper pages and wrapped the images in a PDF shell. The pages look like text but are pictures — no extractor can read them without OCR (optical character recognition), which is a different kind of tool. The giveaway: try selecting text in your PDF viewer. If you can't drag-select a sentence, it's a scan.

Is my PDF uploaded anywhere?
No — never. The file is read directly in your browser by the pdf.js library and every page's text is extracted on your device. Nothing is uploaded, stored, or sent to any server, so it's safe for sensitive documents.
Why does it say 'no embedded text' for my PDF?
That message means your PDF is a scanned (image-only) document: its pages are pictures of text, not actual text. Extracting words from page images requires OCR (optical character recognition), which this tool doesn't do — any extractor that only reads embedded text will come up empty on scans.
Which PDFs can this tool read?
Any standard PDF with embedded text: reports, ebooks, articles, forms, slides saved as PDF. Password-protected PDFs can't be opened, and corrupted or non-standard files will show a clear error instead of hanging.
How do I copy or download the extracted text?
The Copy all button puts every page's text on your clipboard, labeled by page. The Download .txt button saves the same thing as a plain-text file named after your PDF — ready to paste into notes, docs, or study tools.
Does it work offline?
The page needs an internet connection the first time it loads, because the pdf.js extraction engine is fetched from a CDN. After that, the actual extraction happens entirely on your device — your PDF never travels anywhere.
Is there a file size limit?
No hard limit — the extraction runs on your device, so it handles whatever your browser can hold. Very large PDFs (hundreds of pages) simply take longer, with a progress bar showing each page as it's read.

Need help with this tool?

Found a bug, or have a suggestion for the codex? Write to us — a real human reads every message.

[email protected]

More from the codex