More in Edit & Security
About this tool
A closer look at OCR PDF
OCR PDF recognizes text from rendered PDF pages using Tesseract.js, which runs entirely in the browser. Choose a language and a page scope, run OCR, and copy or download the recognized text as a TXT file.
Use it when a document is a scan without selectable text: extracting the text of a scanned contract, making a scan searchable by saving the recognized text, or pulling quotes and data out of image-based pages.
The first OCR run downloads the recognition engine to your browser, which can be slow on a fresh visit; afterwards it runs fully offline. Accuracy depends on scan quality — low-resolution, blurred, handwritten, rotated or noisy pages recognize less reliably. Supported languages are English, Tamil, Hindi, Spanish, French, German and Japanese, including combinations.
Supported files
How to use it
- 1Upload PDF or image
- 2Choose language
- 3Copy or download recognized text
What you can use it for
- Extracting text from scanned documents.
- Making scanned PDFs searchable by saving the OCR text.
- Pulling quotes or data from image-based pages.
- Preparing scanned text for reuse in Word or the AI assistant.
What to keep in mind
- OCR accuracy is lower for low-resolution, blurred, handwritten, rotated or noisy pages.
- The first use downloads the OCR engine locally and can be slow.
- Language options: English, Tamil, Hindi, Spanish, French, German, Japanese, and combinations.
- OCR reads text as image content — it cannot interpret tables or page layout; it outputs plain text.
Troubleshooting
Recognized text has many errors
Re-scan at higher resolution and make sure pages are straight. Rotated pages recognize poorly — fix rotation with the Rotate PDF tool before OCR.
OCR never finishes on a new device
The first run downloads the engine and can take a while. Leave the tab open until the progress indicator completes; later runs are much faster.
Text comes out as one long run
OCR outputs plain recognized text — paragraphs and layout are not reconstructed. Review the TXT or paste it into the AI assistant for structure.
Related tools
Frequently asked questions
Which languages are supported?
English, Tamil, Hindi, Spanish, French, German and Japanese, including combinations.
Why is the first OCR run slow?
The OCR engine is downloaded to your browser on first use, then runs fully offline.
What is the output?
Recognized text that you can copy, or download as a TXT file.
Which scans OCR best?
Sharp, straight, high-resolution scans with plain printed text. Handwriting, rotation and noise all reduce accuracy.
Does my PDF get uploaded for recognition?
No — OCR runs with Tesseract.js inside your browser; the file never leaves your device.