Extract the text from a PDF

Get the text of a PDF as plain text, with lines and paragraphs kept.

Drop a PDF here or click to choose files

Copying from a PDF reader usually gives you a mess: columns interleaved, hyphens everywhere, line breaks in the middle of sentences. That’s because a PDF stores fragments of text with coordinates, not sentences.

This reads those coordinates back and rebuilds the lines — fragments on the same line are joined, and a bigger vertical gap becomes a paragraph break.

If the PDF turns out to be a scan, there is no text to extract at all, and the tool says so and points you at OCR instead.

How to use it

  1. Drop in your PDF.
  2. Wait a moment while the pages are read.
  3. Copy the text, or download it as a .txt file.

Frequently asked questions

Nothing came out — why?

Your PDF is almost certainly a scan: pictures of pages with no text inside. Our image-to-text tool reads those with OCR.

Are tables and columns preserved?

Roughly. Lines are rebuilt in reading order, but a complex multi-column layout or a table will come out as plain lines. For data, exporting from the source document beats extracting from a PDF.

Does it handle accents and other alphabets?

Yes, as long as the PDF embeds the information needed to map its glyphs back to characters, which nearly all modern PDFs do.

Is my document uploaded?

No. It’s parsed in your browser, which matters given how often PDFs are contracts and invoices.

More free tools

All pdf →