OCR PDF
Make scanned PDFs searchable. OCR recognises the text on every page and adds a real, selectable text layer — without changing how your document looks. Free, fast and secure.
or drop a scanned .pdf here
Adds a real text layer so your scanned PDF becomes searchable, selectable and copy-able.
or import from
By uploading files, you agree to our terms of use
Related tools
Make scanned PDFs searchable and selectable by adding an invisible text layer with OCR. oMyPDF's OCR PDF tool uses OCRmyPDF and Tesseract to recognize text in over a dozen languages and embed it invisibly into your PDF — so you can search, copy, and index scanned documents just like any native-text PDF. Processing runs on our servers for the highest recognition accuracy.
How to OCR a PDF
- 1
Upload
Select your scanned PDF, or drag and drop it in.
- 2
Choose language
Pick the document's language and run OCR — text is recognised on every page.
- 3
Download
Download a searchable PDF you can select, copy and search in any reader.
Why a scan is not a document
A scanner produces a photograph. The resulting PDF looks like a document and behaves like a picture: you can see the words, and your computer cannot, because the file contains a grid of pixels rather than characters. Nothing can search it, no text can be copied from it, no screen reader can read it aloud, and no converter can turn it into an editable file.
That gap matters most in exactly the archives where it is most common. A filing cabinet scanned to PDF is not searchable, which means finding one invoice among four thousand is a manual job. Contracts scanned for the record cannot be checked for a clause. Research papers cannot be quoted from without retyping.
Optical character recognition closes it. The page image is analysed, the shapes are recognised as characters, and the resulting text is written into the PDF as an invisible layer positioned exactly over the words it came from. The page looks identical and has become a document.
What determines accuracy
Resolution first. Around 300 DPI is the sweet spot for text — enough detail to distinguish similar letterforms without producing needlessly large files. Below about 200 DPI accuracy falls away quickly, because the fine strokes that separate one character from another stop being captured at all.
Then contrast and geometry. Clean black text on white recognises far better than grey text on a tinted background or a photograph taken at an angle. Skew is handled automatically, as is page rotation, and both help materially — recognition works line by line, so a page tilted a few degrees is much harder than the same page straightened.
The content itself sets the ceiling. Ordinary prose in a standard typeface is the easy case. Dense tables, multiple columns, decorative fonts, mathematical notation and text over images are all harder. Handwriting is a genuinely different problem and standard OCR does not solve it, however good the scan.
Choosing the language
Recognition uses far more than letter shapes. Each language model brings a character set, expected letter combinations and word patterns, all of which resolve the ambiguities that shape alone cannot — the difference between rn and m, between 1 and l, between O and 0 depends on what would make sense in the surrounding word.
Choosing the wrong language therefore degrades every line, not just the accented characters. A German document processed as English loses the umlauts and misreads the long compounds; a Russian or Chinese document processed as English produces nothing usable at all, since the characters are not in the model.
For genuinely bilingual documents, pick the language that carries most of the body text. Names, addresses and short quotations in another language usually survive; a document that is half and half will lose accuracy whichever way you choose.
Where OCR sits in a workflow
Order matters more than it might seem. OCR before compressing, because compression resamples the page images and softer images recognise less reliably — and the text layer itself adds almost nothing to file size, so nothing is lost by doing it in that sequence.
OCR before converting, too. PDF to Word, PDF to Excel and PDF to Markdown all need text to work with; given only images they can only hand back images. Adding the text layer first is what makes an editable result possible at all.
And OCR after redacting rather than before, on any document that needs both. Redaction flattens the pages it touches, which destroys any text layer on them — running OCR afterwards restores searchability while reading only what is now visible, so the redacted content stays gone.
OCR PDF features
Searchable & selectable
Adds an invisible text layer so you can find, select and copy text from scans.
Many languages
Recognise text in English, French, German, Spanish, Chinese, Japanese, Arabic and more.
Looks identical
Your page images are untouched — only a hidden, accurate text layer is added underneath.
Security & privacy
OCR needs a recognition engine far larger than a browser tab can carry, so this runs on our servers. Your PDF travels over an encrypted connection, is processed in an isolated working directory, and both the upload and the result are erased within the hour. The documents that need OCR are overwhelmingly scans of things that matter — contracts, medical records, historical files, personal paperwork — so nothing is archived, indexed, used to train anything, or read by a person.
Frequently asked questions
Is the oMyPDF OCR tool free?
Yes. The Free plan includes 10 OCR conversions a day on files up to 100 MB, with no watermark and no sign-up. Need more? Plus and Pro lift the daily limit.
What does OCR do?
OCR (optical character recognition) reads the text inside scanned page images and adds it back as a real, invisible text layer — so the PDF becomes searchable, selectable and copy-able while still looking exactly the same.
Which languages are supported?
English, French, German, Spanish, Italian, Portuguese, Dutch, Russian, Chinese (Simplified), Japanese, Korean, Arabic and Hindi. Choose the main language of your document for the best accuracy.
Will it change how my PDF looks?
No. The original page images are kept exactly as they are. OCR only adds a hidden text layer underneath, plus optional automatic deskew and rotation for cleaner results.
Can I then edit or convert it?
Yes. Once a PDF is searchable, you can copy its text or run it through PDF to Word, PDF to Excel or PDF to Markdown to get an editable file.
Are my files safe?
Your PDF is processed on our server over an encrypted (HTTPS) connection, OCR'd immediately, and the file is deleted right after.
How do I know whether my PDF needs OCR?
Try to select a word. If the cursor sweeps across the page without highlighting anything, or a search for a word you can plainly see finds nothing, the page is an image and there is no text to find. If text highlights normally, the document already has a text layer and OCR would add nothing.
What accuracy should I expect?
A clean, straight scan at 300 DPI in a common typeface typically lands above 98% on ordinary prose. Accuracy falls with skew, low contrast, unusual fonts, heavy formatting and low resolution, and handwriting is a different problem that standard OCR does not solve. As a rule of thumb, a page that is hard for you to read will be hard for OCR too.
Why does choosing the right language matter so much?
Recognition is not purely shape matching — it uses each language's character set and word patterns to resolve ambiguity, which is how it tells rn from m or 1 from l. Running a French document as English costs accuracy on every accented character and every word the English model does not expect.
Will the page still look the same?
Yes. The original image is kept exactly as it is and the recognised text is written invisibly behind it, aligned to the words it came from. You see your scan; your software sees text. Deskew and rotation correction are applied where they help recognition.
Does OCR make a scan editable in Word?
It is the prerequisite. Without a text layer, PDF to Word can only hand you a picture per page; with one, it can reconstruct real paragraphs and tables. So the sequence for an editable scan is OCR first, convert second.
Should I OCR before or after compressing?
Before. Compression resamples the page images, and softer images recognise less accurately. OCR the good-quality scan first, then compress — the text layer costs almost nothing in file size and survives compression intact.