Skip to content
GivenTool

OCR a scanned PDF

Make scanned PDFs searchable: Tesseract reads English and Arabic text and adds an invisible text layer.

Loading tool…

A scanned PDF is a stack of pictures: you cannot search it, select a sentence or copy a number. OCR (optical character recognition) reads the text in those pictures. Each page is rendered at 200 or 300 dpi and read by Tesseract, the open-source OCR engine, running in your browser. The words it finds are laid over the page as an invisible text layer, so the PDF looks exactly as before but can be searched, selected and copied; the text is also saved as a .txt file.

English and Arabic are supported, separately or together. Pages that already contain text can be skipped, so mixed documents are handled correctly. Progress is shown page by page, and you can cancel at any time.

OCR is never perfect. Clean, straight scans of printed text at 300 dpi give excellent results; handwriting, low resolution, skewed phone photos, stamps and decorative fonts give poor ones. Check important names and figures against the page. The first run downloads the engine and language data (about 7-9 MB) from this site.

How to use it

  1. Choose a scanned PDF.
  2. Pick the document language and resolution.
  3. Click Run OCR and follow the page-by-page progress.
  4. Download the searchable PDF and the .txt file.

Frequently asked questions

Are my files uploaded?

No. The first time you use the tool your browser downloads the processing engine from this site, like any other part of the page; your PDF itself is read and processed on your device and never sent anywhere.

Does OCR change how my PDF looks?

No. Your pages stay exactly as they are; an invisible text layer is added on top, aligned with the words, so selecting and searching works.

How accurate is it?

For clean printed text at 300 dpi, typically very accurate. Accuracy drops with blur, skew, low resolution, small or decorative fonts and handwriting, which Tesseract does not read reliably.

Can it read Arabic?

Yes, with the Arabic model. Words are recognised correctly in most clean scans; the order in which copied Arabic text is pasted can vary between PDF readers.

How long does it take?

Roughly 2-10 seconds per page on a modern laptop at 300 dpi, longer on phones. Use 200 dpi for faster, slightly less accurate results.

What if my PDF is password-protected?

The tool will say so. If it is your file and you know the password, remove the protection with Unlock PDF first, then use the unlocked copy here.