OCR a scanned PDF
Make scanned PDFs searchable: Tesseract reads English and Arabic text and adds an invisible text layer.
Loading tool…
A scanned PDF is a stack of pictures: you cannot search it, select a sentence or copy a number. OCR (optical character recognition) reads the text in those pictures. Each page is rendered at 200 or 300 dpi and read by Tesseract, the open-source OCR engine, running in your browser. The words it finds are laid over the page as an invisible text layer, so the PDF looks exactly as before but can be searched, selected and copied; the text is also saved as a .txt file.
English and Arabic are supported, separately or together. Pages that already contain text can be skipped, so mixed documents are handled correctly. Progress is shown page by page, and you can cancel at any time.
OCR is never perfect. Clean, straight scans of printed text at 300 dpi give excellent results; handwriting, low resolution, skewed phone photos, stamps and decorative fonts give poor ones. Check important names and figures against the page. The first run downloads the engine and language data (about 7-9 MB) from this site.
How to use it
- Choose a scanned PDF.
- Pick the document language and resolution.
- Click Run OCR and follow the page-by-page progress.
- Download the searchable PDF and the .txt file.
Frequently asked questions
Are my files uploaded?
No. The first time you use the tool your browser downloads the processing engine from this site, like any other part of the page; your PDF itself is read and processed on your device and never sent anywhere.
Does OCR change how my PDF looks?
No. Your pages stay exactly as they are; an invisible text layer is added on top, aligned with the words, so selecting and searching works.
How accurate is it?
For clean printed text at 300 dpi, typically very accurate. Accuracy drops with blur, skew, low resolution, small or decorative fonts and handwriting, which Tesseract does not read reliably.
Can it read Arabic?
Yes, with the Arabic model. Words are recognised correctly in most clean scans; the order in which copied Arabic text is pasted can vary between PDF readers.
How long does it take?
Roughly 2-10 seconds per page on a modern laptop at 300 dpi, longer on phones. Use 200 dpi for faster, slightly less accurate results.
What if my PDF is password-protected?
The tool will say so. If it is your file and you know the password, remove the protection with Unlock PDF first, then use the unlocked copy here.
Related tools
- Extract text from a PDFExtract the text of a PDF in reading order and download it as a .txt file or copy it.
- Convert PDF to WordTurn a text-based PDF into an editable Word document with paragraphs and headings rebuilt.
- Scan documents to PDFPhotograph pages with your camera, clean them up to look like scans and save them as one PDF.
- Compress a PDF to a smaller file sizeShrink scanned PDFs to a target size (such as 200 KB) or by level, on your device. The result is image-based.
- Compare two PDF filesFind what changed between two versions of a PDF: added and removed words, shown page by page.