How to recognise text in a scanned PDF?
- 1.
Load the PDF
Drop the file or choose it. If it is password-protected, enter the password. It never leaves your device.
- 2.
Set languages and range
Select the document languages and optionally a page range, e.g. “1-3, 5”.
- 3.
Run OCR
Click “Recognise text”. On first use the OCR engine and language data are downloaded (about 7 MB for Polish and English). Recognition time depends on your device, and you can cancel at any time.
- 4.
Download
Copy the text, download it as .txt or download a searchable PDF: the original pages with an invisible text layer.
Scanned PDF or text PDF?
Try selecting text in your PDF viewer. If you can, the file has a text layer and OCR is unnecessary: PDF to text is faster and makes no recognition errors. If the whole page selects like an image, it is a scan and needs OCR. The tool checks up to 5 pages and warns you when the document already has text.
For a document containing both scans and digital text, choose only the scanned pages in the range field. Review names, dates and amounts against the original before using recognised text in another document.
How does in-browser PDF OCR work?
Each selected page is rendered at 300 DPI (an A4 page is about 2480 × 3508 px) and recognised with the Tesseract engine (Apache-2.0) using LSTM models for the selected languages. Your file is not uploaded. Only the OCR engine (about 1.4 MB) and the language data (Polish 2.6 MB, English 2.9 MB, German 1.3 MB, Ukrainian 2.1 MB) are downloaded from the jsDelivr CDN. On later visits the browser usually loads them from its cache.
What is a searchable PDF?
A searchable PDF looks like the original scan but has an invisible text layer under the image, placed where the words were recognised. Search (Ctrl+F), text selection and indexing then work, while the look of the pages stays the same. The option is not available for password-protected files, which can be read and recognised but not re-saved in the browser.
Processing time depends on your device, the number of pages and the number of selected languages. For long documents choose a page range. Best results come from straight, high-contrast 300 DPI scans; for faded print enable image enhancement.
Confidence score and typical mistakes
For every page the tool shows the engine’s average confidence. Above 80% the text usually needs only minor corrections; below 60% expect many errors. The most common causes are skewed or blurry scans, very small print, coloured backgrounds and a missing language. Without Polish selected, letters such as ł or ę are read as t and e. Single photos and screenshots are easier to handle with the image to text tool.
Frequently asked questions
Is PDF OCR free?
+
Is my document uploaded to a server?
+
Can I OCR a password-protected PDF?
+
How do I make a searchable PDF from a scan?
+
What if my PDF already contains text?
+
Guides using this tool
How to prepare a PDF for email
Your PDF is too big to attach, or someone asked for a smaller copy. Find out what is taking up the space, remove pages nobody needs, compress, and look at the result before you send it.PDFHow to scan a document with your phone to PDF
A phone and decent light are enough to scan a document. The photo still needs to be straightened, cropped to the page, and saved as a PDF that the recipient can actually read.PDFWhy you can't select text in a PDF
You drag across a PDF and the whole page gets selected, pasted text comes out as gibberish, or Ctrl+F finds nothing. Each symptom has a different cause and a different fix.Updated: