PDF to Text Extractor

Extract text from a PDF without copying page by page. This tool reads the document's text layer, preserves the reading order, and gives you the whole thing as plain text you can copy or download.

Your files never leave this device

Extract the text One file at a time ยท 100 MB

Drag & drop a PDF here

Works on documents with a real text layer. Scanned pages hold images, not text, and cannot be read.

No files selected yet

    Formatting

    Turn line breaks off to get flowing paragraphs instead of one line per printed line.

    How to extract text from a PDF

    1. Drop your PDF onto the box above, or press Browse files to select it.
    2. Decide whether to keep the original line breaks or let the text flow into paragraphs.
    3. Tick the page marker option if you need to know which page each passage came from.
    4. Press Extract text and wait while the pages are read.
    5. Review the text in the box below, then download it as a .txt file.

    What this extractor does

    • Reads the whole document in one pass instead of page-by-page copying.
    • Follows the text layer's reading order rather than raw drawing order.
    • Optionally preserves line breaks, or runs lines together into flowing paragraphs.
    • Can mark where each page begins, which matters for citation and reference.
    • Reports the word count so you can sanity-check the result immediately.
    • Says plainly when a document is a scan with no text to extract.

    Two kinds of PDF, and only one of them has text

    This distinction explains almost every question about text extraction. A PDF created from a word processor, a web page or design software contains the actual characters, positioned on the page with font information attached. Open it in a reader, drag across a paragraph, and the text highlights. That is the text layer, and this tool reads it directly.

    A PDF created by a scanner or a camera contains photographs of pages. To a computer there is no text in it at all, only pixels that happen to look like letters to a human. Dragging across a line highlights nothing. No amount of extraction can produce words from that, because there are none stored.

    How to tell which you have

    • Open the PDF and try to select a sentence. If it highlights, there is a text layer.
    • Use the reader's search function. If searching for a word you can see finds nothing, the page is an image.
    • Zoom in a long way. Text stays crisp; a scan becomes visibly pixelated.

    Turning a scan into text needs optical character recognition, which analyses the shapes of the letters and guesses what they are. This tool does not do OCR, and it says so rather than returning an empty file and leaving you to work out why. If that is what you have, export the pages with PDF to PNG and run them through the image to text converter, which does.

    Line breaks are a real decision

    PDFs store text positioned on a page, not as flowing paragraphs. A printed line ends because the layout ran out of width, not because a sentence finished. When line breaks are preserved you get text that mirrors the printed page, which is right for poetry, code, tables and anything where the layout carries meaning.

    Turn the option off and lines are joined into continuous paragraphs, which is what you want before pasting into a word processor, feeding text into a translator, or running a word count. Reflowing later is tedious; choosing correctly here takes one click. If the extracted text still needs tidying, the line break remover and text cleaner handle the rest.

    Bengali, Hindi and other Indic scripts

    These scripts need a step that Latin text does not. Several vowel signs are written to the left of their consonant but must be stored after it, and a PDF hands back its characters in the order they were drawn. Left alone, the Bengali for Bangladesh comes out as a jumble with the vowel sign sitting in front of the wrong letter, and the same happens to almost every word on the page.

    This tool puts them back. The vowel is moved past the whole consonant cluster rather than just the first letter, so conjuncts survive. On a real Bangladesh Railway ticket that corrected sixty one of the sixty seven misplaced vowel signs in the document.

    The remaining few cannot be repaired by anyone, and it is worth explaining why. Some PDFs map their ligature glyphs to the private use area of Unicode, which is a way of saying that the file itself declares those shapes have no letter behind them. The information was discarded when the PDF was made. Where that happens the extractor reports how many characters were affected rather than quietly handing you text with holes in it.

    What extraction cannot preserve

    Plain text is plain. Bold and italic vanish, along with font sizes, colours, headings and indentation. Tables are the most painful loss: their cells become a stream of words, because a PDF stores a table as text at coordinates rather than as rows and columns. Multi-column layouts can interleave for the same reason, although the reading order usually holds up well in practice.

    Images and figures are not text and do not appear. If you need those, the image extractor pulls them out separately, and PDF to PNG captures whole pages as pictures.

    Useful things to do with the result

    Once the text is out, it becomes searchable, countable and editable. Run it through the word counter to check a length requirement, the word frequency counter to see what a long report actually emphasises, or the case converter to normalise a document typed in capitals.

    Extraction runs entirely in your browser, so contracts, reports and research papers are never uploaded anywhere.

    Frequently asked questions

    Why did my PDF return no text?

    It is almost certainly a scan, which stores pages as images rather than characters. Try selecting text in a PDF reader. If nothing highlights, there is no text layer to extract.

    Does this tool do OCR on scanned documents?

    No. It reads the existing text layer only. For a scan, export the pages as images with PDF to PNG and use the image to text converter, which does run optical character recognition.

    Should I keep the line breaks?

    Keep them for code, poetry, tables and anything where layout matters. Turn them off before pasting into a word processor or translator, so the text flows as paragraphs.

    Why does Bengali or Hindi text come out scrambled elsewhere?

    Because several Indic vowel signs are drawn to the left of their consonant but must be stored after it, and PDFs return characters in drawing order. This tool reorders them, moving each vowel past the whole consonant cluster.

    Some characters are still missing from my Bengali PDF. Why?

    Certain PDFs map ligature glyphs to Unicode's private use area, which means the file itself records no letter for those shapes. That information was lost when the PDF was created, so no tool can recover it. The count of affected characters is reported with the result.

    Why did my table come out jumbled?

    A PDF stores a table as text placed at coordinates, not as rows and columns. Extraction produces a stream of words in reading order, so table structure cannot survive.

    Is my document uploaded to extract the text?

    No. The whole document is read inside your browser, so contracts, reports and unpublished work never leave your device.