Guide
Run OCR on a scanned PDF in the browser
Use Tesseract in your tab to add a text layer to image-only scans so you can search or convert to Word afterward.
How to OCR a PDF
Phone photos of paper have no text layer. OCR guesses characters from pixels. Tesseract runs in WebAssembly in your browser, which is why the first run may download language data and why it is slower than a native app.
Handwriting, stamps, and tiny footnotes fail often. Clean scans of printed Latin text work better. This is not a human typist.
Your scan does not go to a cloud OCR API operated by SmartPDFConvert. The browser does the work, which is also why a 50-page scan can heat up a laptop.
- Open OCR PDF and add a scanned PDF. Pages that are already digital text do not need this step.
- Click Run OCR and leave the tab open. Optical recognition is slow, especially on phones.
- Download the searchable PDF when progress finishes.
- Search for a word you can see on the scan. If the scan is skewed or low contrast, expect errors.
Browser OCR
If search still finds nothing, the pages may be too dark or not actually images of text. Try a better scan rather than repeating OCR.
- Adds searchable text using Tesseract.js in the page.
- Helps PDF to Word and PDF to Excel on former scans.
- Accuracy depends on scan quality and language.
- Can be slow; keep the tab in the foreground.
- No SmartPDFConvert OCR server.
Privacy for OCR
Pixels are processed on your device. OCR can still misread a name into the text layer; check before you share.
Tesseract language files may be loaded as scripts for the page. Document images are not posted to our conversion API. See the Privacy Policy.
Frequently asked questions
The engine’s default in this tool is aimed at common printed text. Specialty languages may need a desktop OCR pack.
The visible scan stays. A text layer is added for search and copy.
Each page is expensive. Wait, or split the PDF and OCR a few pages at a time.
