PDF Cubby
Sign in
Conversion

How to Convert a Scanned PDF to Word (Without Getting Gibberish)

PDF Cubby · July 21, 2026 · 6 min read

You drop a scanned document into a PDF-to-Word converter, wait, and get back a Word file with nothing in it — or a page of nonsense characters. The converter is not broken and neither is your file. A scan simply does not contain any text for it to convert, and almost nothing tells you that before you start.

Why a scanned PDF converts to an empty Word file

There are two completely different things people call a PDF. One is born digital — exported from Word, a browser, or an accounting package. Every character in it is stored as an actual character, which is why you can select a sentence and copy it.

The other is a scan: a photograph of a page, wrapped in a PDF. It looks identical on screen, but the words are pixels. There are no characters in the file at all. When a converter opens it looking for text to rebuild in Word, it finds nothing — so it hands back an empty document, or it guesses and produces gibberish.

A thirty-second test: open the PDF and try to select a sentence with your cursor. If nothing highlights, it is a scan, and no converter will get text out of it as-is.

The step that is missing: OCR

OCR — optical character recognition — reads the picture of the page and works out which shapes are which letters, then writes those letters back into the file as real, selectable text. The page still looks exactly the same. What changes is that there is now a text layer underneath it.

Once that layer exists, converting to Word behaves normally, because there is finally something to convert.

How to convert a scanned PDF to Word properly

  1. Open OCR and add your scanned PDF. Pick the language the document is written in — accuracy depends on it more than most people expect.
  2. Run it, and download the result. It will look unchanged, which is correct. Try selecting a line of text: this time it should highlight.
  3. Take that OCR'd file to PDF to Word.
  4. Download your .docx and open it. Headings, paragraphs and straightforward tables should be there and editable.

What to expect from the result

OCR is very good, not perfect. Clean printed text at a reasonable resolution converts extremely well. Handwriting, faint fax copies, heavy background patterns and unusual fonts are where mistakes creep in — usually single characters rather than whole words.

Scan quality is the single biggest factor. If you still have the paper original and the first result disappoints, rescanning at 300 dpi will improve things more than any change of tool.

If your file is a photo rather than a scan

Pictures of documents taken on a phone work the same way, with one extra step: turn the images into a PDF first with Image to PDF, then OCR that. Photograph the page flat, in even light, with the whole page in frame — shadows and angles cost more accuracy than a slightly lower megapixel count.

A note on file size

OCR adds a text layer without touching your original image, so the file grows only slightly — a few percent. If the scan was already too large to email, Compress it after OCR rather than before, so the text layer is built from the sharper original.

What OCR is actually doing

It helps to know roughly what is happening, because it explains most of the surprises. The software goes across the page looking for connected regions of dark pixels that sit apart from their neighbours, and treats each one as a candidate character. It then compares each shape against what it knows about letterforms in the language you selected, and picks the most likely match.

Crucially, it also uses context. A shape that could be a capital I or a lowercase l or the digit 1 gets decided partly by what surrounds it — which is why a word in the middle of a sentence is read more reliably than a lone character in a table cell, and why an account number sitting on its own is one of the riskier things on the page.

This is also why language selection matters so much. Choosing the wrong language does not just affect accented characters; it changes which letter combinations the engine considers plausible, and that shifts guesses across the whole document.

How different documents tend to fare

Getting tables to survive the trip

Tables are where converted documents most often disappoint, and it is worth understanding why. A table in a scan is not stored as a table — it is lines and text that happen to be arranged in a grid. Rebuilding it means inferring the structure from the position of those lines, and that inference is easier for some tables than others.

Tables with visible ruled borders on every side convert most reliably, because the boundaries are unambiguous. Tables that rely on whitespace alone, or on shading rather than lines, are much harder — the converter has to guess where one column ends and the next begins, and merged or nested cells frequently end up split or misaligned.

If the table is the point of the document, it is worth checking the result cell by cell before reusing it. And if you only need the numbers rather than the layout, pulling the content out with PDF to text and rebuilding the table yourself is often faster than fixing a mangled one.

Check the converted document before you rely on it

  1. Compare the page count. If the Word file has fewer pages than the PDF, something was skipped.
  2. Spot-check the numbers. Account numbers, totals, dates and reference codes. These are where an OCR error does real damage and where proofreading by eye is least likely to catch it.
  3. Read the first and last paragraph of each section. Reading-order problems show up at section boundaries before anywhere else.
  4. Search the Word file for a distinctive phrase you know is in the original. If it is missing, that region did not convert.
  5. Check anything that was in a different font or size — headers, footers, stamps and marginal notes are the most commonly dropped elements.

When you do not need Word at all

Converting to Word is the right answer when you genuinely need to edit the document: rewriting clauses, updating figures, reusing a section as the basis for something new. It is a heavier operation than people assume, and it is not always what the job requires.

If you only want to copy a few paragraphs out, run OCR and then simply select and copy from the PDF — no conversion needed. If you want the raw wording with no formatting at all, PDF to text is faster and cleaner. And if you only need the document to be findable rather than editable, OCR on its own is enough; the file stays a PDF and stays exactly as it looked.

Keep the original

Whatever you do next, hold on to the untouched scan. A converted Word document is an interpretation of the original, not a replacement for it — it carries whatever mistakes the recognition made, and those mistakes are invisible once the source is gone. If the document has any legal or financial weight, the scan is the record and the .docx is a working copy.

This matters most with the documents people most often convert: contracts, statements, payslips, invoices. Convert freely, edit freely, but file the original somewhere you will still be able to find it in a year. Nothing you upload is kept on our side once your download is ready, so the only copy that survives is the one you keep.

Try it now

PDF to Word is free, needs no sign-up, and your file is never stored.

Open PDF to Word

FAQ

Why is my converted Word document empty?
Because the PDF is a scan — a picture of a page rather than text. There are no characters in the file for the converter to rebuild, so it returns nothing. Run OCR first to create a real text layer, then convert.
Can I convert a scanned PDF to Word for free?
Yes. OCR and PDF to Word each give you five free uses a day with no payment, which covers occasional documents. Pro removes the daily limit if you work with scans regularly.
Will the formatting survive?
Headings, paragraphs and straightforward tables usually come across in place. Complex multi-column layouts and nested tables are where a converted document is most likely to need tidying.
Does OCR change how my scan looks?
No. The text layer is invisible and sits underneath the existing image, so the page looks exactly as it did. Your original scan quality is preserved, not re-rendered at a lower resolution.
Which languages are supported?
Thirteen, including English, French, Spanish, German, Portuguese, Italian, Dutch, Polish, Chinese, Japanese, Arabic, Russian and Hindi. Choosing the right one noticeably improves accuracy.
What happens to my document?
It is processed and discarded the instant your download is ready. It is never stored, never queued, and never read by a person.