PDF Cubby
Sign in
OCR

How to Make a Scanned PDF Searchable

PDF Cubby · July 21, 2026 · 5 min read

You have a scanned contract, lease or stack of invoices, you press Ctrl+F to find one name, and the search returns nothing — even though the word is clearly on the screen in front of you. Nothing is broken. The page is a picture, and search cannot read pictures.

Why search finds nothing

A scanner produces an image of a page. That image gets wrapped in a PDF so it is easy to share and print, but the words inside it are just areas of dark pixels. Your eyes read them effortlessly; the software has no idea they are words at all.

This is why a scanned document cannot be searched, cannot be copied from, and returns nothing to a text converter. The fix is not a better PDF reader — it is adding the missing text.

What making a PDF searchable actually does

OCR examines the image, identifies the shapes as characters, and writes those characters into the file as a text layer positioned exactly over the words they match. The page still displays the original scan, so it looks unchanged. Underneath, there is now real text that search, selection and copying can all use.

This is why a searchable scan looks identical to the original. Anything that visibly re-renders or softens your pages has done more than add a text layer, and has usually lost quality doing it.

How to make a scanned PDF searchable

  1. Open OCR and add your scanned PDF.
  2. Choose the document's language. This matters more than people expect — the wrong language guesses the wrong characters.
  3. Run it. Longer documents take longer; a rough guide is a couple of seconds per page at normal scan resolution.
  4. Download the result and test it — see below.

How to check it actually worked

Do not take the word “done” for it. Two quick checks:

If a word genuinely on the page still cannot be found, the usual cause is scan quality rather than the tool — see below.

When results disappoint

What becomes possible afterwards

A searchable scan stops being a dead end. You can find a clause in a hundred-page lease in seconds, copy a paragraph without retyping it, convert the document to an editable file with PDF to Word, or redact it properly — because Secure Redact works by finding text, and until now there was none to find.

That last point catches people out. Drawing a black box over a name in a scan hides it visually but removes nothing. Redaction can only delete text that exists, which is why OCR has to come first.

What the software is actually doing

Knowing roughly how OCR works explains most of its quirks. It scans the page for connected clusters of dark pixels, treats each cluster as a possible character, and matches the shape against what it knows about letterforms in the language you chose. It then leans on context to settle the ambiguous cases — whether a mark is a capital I, a lowercase l or the digit 1 often depends on the characters either side of it.

That context-dependence has a practical consequence. Words inside ordinary sentences are recognised more reliably than isolated strings, which is why a reference number sitting alone in a box is among the least reliable things on any page, and worth checking by eye.

Why searchable matters beyond Ctrl+F

Finding a word on the page is the obvious benefit, but it is rarely the most valuable one.

Getting the best result from the scan itself

OCR accuracy is decided mostly before the file ever reaches the software. If results disappoint and you still have the paper original, changing how you scan will help more than changing tools.

Searchable PDF or plain text extraction?

These are different jobs and people often reach for the wrong one. Making a PDF searchable keeps the document intact — same pages, same appearance, same signatures and stamps — and adds an invisible text layer underneath. The file remains the document, and remains something you can send to someone else as a faithful copy.

Extracting to plain text throws the presentation away and gives you the words alone. That is ideal for feeding content into something else, and useless if you need the document to still look like the document.

As a rule: if you need to keep, share or archive the document, make it searchable. If you only need what it says, extract the text with PDF to text.

Working through a batch

If you are digitising a stack of paperwork rather than one file, a little order helps. Straighten pages first with Rotate, since skew costs accuracy across every page that follows. Run OCR next, while the images are still at full quality. Compress last, once the text layer already exists — doing it the other way round means recognising characters from a degraded image.

For long documents, OCR time scales with page count and resolution, so a two-hundred-page scan at high resolution is a genuinely slow job rather than a stuck one. Splitting very long documents into sections with Split makes progress visible and means a problem with one section does not cost you the whole run.

A note on archives

If these documents are going into long-term storage, making them searchable is worth doing once, at the point of filing, rather than later when you need something urgently. A drawer of scanned PDFs that cannot be searched is only marginally better than the paper it replaced — you still have to open each file and read it to find anything.

It also protects against a quieter problem: the person who filed the documents knows what is in them, and that knowledge leaves when they do. Text inside the files does not.

Common questions people hit on the first run

If a document still resists after a clean 300 dpi re-scan with the right language selected, it is usually telling you something real: the print is too degraded, or it is handwritten. At that stage retyping the handful of lines you actually need is faster than fighting it.

Try it now

OCR is free, needs no sign-up, and your file is never stored.

Open OCR

FAQ

Why can I not search my scanned PDF?
Because the pages are images. The words are pixels rather than characters, so there is nothing for search to match. OCR adds a real text layer and search then works normally.
Does making a PDF searchable change how it looks?
No. The text layer is invisible and sits beneath the existing page image, so the document looks exactly as it did before, at the same quality.
How much bigger does the file get?
Only slightly — typically a few percent. The added text is tiny compared with the page images, and your original images are left untouched.
How long does it take?
Roughly a couple of seconds per page at normal scan resolution, longer for high-resolution scans. You get a time estimate before it starts.
Is it free?
You get five free uses a day with no payment. Pro removes the daily limit for people who work with scans regularly.
Can it read handwriting?
Not reliably. OCR is built for printed text. Handwritten notes may produce partial or incorrect results.
What happens to my file?
It is processed and discarded the moment your download is ready — never stored, never queued, never seen by a person.