Skip to content
HNarzędzia
en
Categories

PDF to text online - extract and copy text from a PDF

Pull the text out of a PDF in one step: copy it to the clipboard or download it as a plain .txt file. It runs locally in your browser and reads the PDF’s text layer - scanned pages need OCR instead.

  • Free
  • No sign-up
  • Private
  • Runs locally

Drag & drop files here

or click to choose from your device

PDF file · processed locally in your browser

How to extract text from a PDF

  1. 1.

    Open the PDF

    Drop the file onto the area above or choose it from your device. If it is password-protected, enter the password - it is used only on your device.

  2. 2.

    Wait for extraction

    The tool reads the pages one by one and shows progress. A few hundred pages usually take a few seconds.

  3. 3.

    Adjust the output

    “Join broken lines” is on by default and rebuilds paragraphs. Turn on “Add page headers” if you need to know which page each passage came from.

  4. 4.

    Copy or download

    Click “Copy text” to paste it anywhere, or “Download .txt” to save a UTF-8 file named after your PDF, for example name.txt.

Where the text comes from - and when there is none

PDFs created from Word, Google Docs, accounting software or a browser’s “Save as PDF” contain a text layer: the actual characters with their positions on the page. This tool reads that layer with pdf.js, so the result is exact - no recognition errors, no guessing.

Scans and photos of documents are different. A page from a multifunction printer or a phone scanning app is just an image, and there is no text to extract. If the whole file averages fewer than about 15 characters per page, the tool reports that it has no text layer; if only some pages are empty (say, a signed and scanned last page), it warns that those pages may need OCR. A quick test: if you cannot select words with the mouse in your PDF viewer, it is almost certainly a scan. For those files use OCR PDF, or Image to Text for single photos.

Joining broken lines into paragraphs

A PDF stores every line separately, so copying a paragraph usually gives you a ladder of short lines. Join broken lines fixes that: lines inside a paragraph are merged with a space, while a line ending with a period, question mark, exclamation mark, colon or semicolon closes the sentence and the next line starts on a new row. Words split by a hyphen at the end of a line are glued back together, so “docu-” + “ment” becomes “document”. Bullet points, dashes and numbered items such as “1.” or “2)” keep their own lines.

For tables, forms, invoices, poetry or source code, turn the option off - the line layout then stays exactly as the PDF stores it, with long runs of blank lines trimmed.

Page headers, statistics and what to do next

Add page headers inserts a line like --- Page 3 --- before each page’s text. It helps when you quote a document and need a page reference, or when you search for the page a clause was on. Without headers, empty pages are skipped and pages are separated by a blank line.

Below the output you see the number of pages, words and characters - a 12-page contract is typically 4,000-5,000 words. Need only a few pages? Cut them out first with Split PDF, then load the smaller file here. If you want images rather than text, use Extract images from PDF.

Privacy and limitations

  • Extraction runs entirely in your browser; the PDF is never uploaded, which makes it safe for contracts, payslips and medical letters.
  • The .txt file is saved in UTF-8, so accented letters, symbols and non-Latin scripts are preserved and open correctly in Notepad, TextEdit, VS Code or Word.
  • Multi-column layouts, tables and footnotes can come out in a different order than you see on the page, because PDF stores positioned text fragments, not paragraphs or reading order.
  • Formatting such as bold, fonts, links and images is not kept - the output is plain text.
  • Some PDFs have broken font mappings and produce gibberish instead of letters; in that case OCR is the only way to get usable text.

Frequently asked questions

Why can’t I copy text from my PDF?

+
Most often the file is a scan - a picture of a page with no text layer - so it needs OCR. Less often the fonts have broken character mappings, which also produce unreadable output.

Does this tool do OCR on scanned PDFs?

+
No. It reads only the existing text layer and tells you when there is none. For scanned documents use the OCR PDF tool, which recognises text from page images.

What format is the downloaded file?

+
A plain .txt file in UTF-8 encoding, named after your PDF. It keeps accented and special characters and opens in any modern text editor.

How do I remove line breaks from text copied out of a PDF?

+
Keep “Join broken lines” switched on. It merges lines that belong to one paragraph and rejoins words hyphenated at the end of a line.

Can I extract text from a password-protected PDF?

+
Yes, if you know the password to open it. A password field appears after you load the file, and the password is used only locally.

Is my PDF uploaded to a server?

+
No. The text is extracted in your browser and the document never leaves your device.

Updated: