How to extract text from an image?
- 1.
Choose languages
Select the languages of the text - Polish and English by default; German and Ukrainian are also available.
- 2.
Add an image
Drop a photo or scan, choose a file or paste a screenshot with Ctrl+V.
- 3.
Wait for recognition
On first use the browser downloads the OCR engine and language data (about 7 MB for Polish + English). An A4 page usually takes 3-10 seconds.
- 4.
Copy or download
Correct any mistakes in the text box, then click “Copy text” or “Download .txt”.
How does in-browser OCR work?
OCR (optical character recognition) turns characters in an image into editable text. The tool uses the Tesseract engine (Apache-2.0) compiled to WebAssembly, with an LSTM neural network trained separately for each language.
Your image is not uploaded. Only the OCR engine (about 1.4 MB) and the selected language data are downloaded from the jsDelivr CDN, once, and then cached by your browser.
After recognition, review the result before using it. OCR can confuse similar characters such as O and 0, especially in low-quality scans, unusual fonts or small labels.
Download size
| Component | One-time size |
|---|---|
| OCR engine (Tesseract 5, WebAssembly) | about 1.4 MB |
| Polish | about 2.6 MB |
| English | about 2.9 MB |
| German | about 1.3 MB |
| Ukrainian | about 2.1 MB |
Tips for better accuracy
- Shoot straight on - avoid perspective and shadows.
- Resolution - letters should be at least 20 px tall; small screenshots are enlarged ×2 automatically.
- Contrast - for grey paper, receipts and faint print turn on “Enhance the image” (greyscale and contrast stretch).
- Correct language - language data contains the special letters and typical words of each language.
- Handwriting - Tesseract is trained on printed text and handles handwriting poorly.
The confidence value is the engine’s average score for all words. Above 80% the text usually needs only minor fixes; below 60% try a better photo.
Photo, scan or PDF?
This tool accepts a single image: JPG, PNG, WebP or BMP. Multi-page scans saved as PDF can be recognised with the OCR PDF tool, which processes the document page by page and shows progress for each page. If your PDF already has a text layer (you can select text in it), a text extractor is much faster than OCR, because it reads the text directly instead of recognising it from pixels. Typical uses of image OCR: copying text from a screenshot of a video call or presentation, digitising a printed letter, a receipt or a page from a book, and quoting a fragment of an infographic.
Frequently asked questions
Is my image uploaded to a server?
+
How do I copy text from a screenshot?
+
Does it recognise handwriting?
+
Why does the text contain mistakes?
+
Can I extract text from a scanned PDF?
+
Updated: