Convert a scanned PDF to Word with OCR

A scanned PDF holds pictures of pages, so PDF to Word reports zero text and Word's converter returns a document of images. Optical character recognition, OCR, has to read the pixels first. This guide shows how to confirm the PDF is a scan, which OCR routes keep the file on your machine, and how to check the text afterwards, because OCR always misreads some characters.

PDF to Word

How to convert scanned PDF to Word with OCR

  1. Confirm it is a scan

    Open the PDF and try to select a word. If the cursor selects a whole page as one block, or nothing, there is no text layer. PDF to Word reports zero text for such a file.

  2. Run an OCR step

    On a Mac, open a page in Preview and select the text directly. Otherwise use an OCR step in a desktop app or a scanner app that writes a searchable PDF, or Drive's Open with Google Docs if sending the file to Google is acceptable.

  3. Get the text into Word

    Convert the searchable PDF in PDF to Word, paste from Preview, or download the Google Doc as .docx. Apply styles for headings and lists; OCR delivers plain paragraphs.

  4. Proofread against the scan

    Put the scan and the Word file side by side and read every number, name and date. OCR confuses 0 and O, 1 and l, 5 and S.

Why the converter gave you pictures

Every PDF to Word converter reads the text layer. A scanner and a phone camera do not write one; they write a JPEG of each page inside a PDF wrapper. Word's converter finds no text, so it places each page image into the document and stops. The file opens, looks right, and cannot be edited or searched, which is the moment most people go looking for OCR.

PDF to Word in this tab reads the same text layer. On a scan it reports zero text and offers no download, because a .docx of page pictures is not a Word document. The general converter routes, for PDFs that do have text, are in Convert PDF to Word. Nothing there works on a scan until one of the routes below has run.

OCR PDF to Word on your own machine

macOS has OCR built in. Open the scanned PDF in Preview, drag across the text on a page, and Live Text selects the words it has read. Copy, paste into Word, repeat per page. It runs on the Mac and nothing leaves it. Windows has the same capability in PowerToys Text Extractor and in OneNote's Copy Text from Picture, both local, both one page at a time.

For a long scan, an OCR step in a desktop app or a scanner app that writes a searchable PDF in one pass saves hours. The searchable PDF keeps the page image and adds a hidden text layer, and that text layer is what PDF to Word then reads, rebuilds into paragraphs and counts before the download. Any of these keeps the file on your machine, which for a scanned passport, contract or medical record is the point.

Scanned PDF to Word through Google Docs

Upload the PDF to Google Drive, right-click it, choose Open with, then Google Docs. Drive runs OCR on Google's servers and opens a Doc with the recognised text, one paragraph per block it found, with most layout gone. File, Download, Microsoft Word gives you a .docx. It handles many languages and is the shortest route for a document with no privacy weight.

The file leaves your machine, and a long or heavy scan fails or comes back with only its first pages. Both are stated plainly because the route is otherwise tempting for everything.

What PDF to Word does with the OCR output

A searchable PDF from an OCR step converts like any native-text PDF. The tool reads the hidden text runs with their positions, rebuilds paragraphs, writes the .docx, then reads it back and counts the text before it offers the download. The count is the OCR engine's text, so a misread digit is counted as present; the count proves the conversion kept what the OCR found, not that the OCR found the right characters.

That split matters. The conversion step is checked by the tool. The recognition step is checked by you, against the scan, and no count replaces that reading.

Checking OCR output

OCR accuracy on a clean 300 dpi office scan is high but not complete, and the errors cluster in exactly the places that matter: digits, names and dates. A 0 becomes an O, a 1 becomes an l, an rn becomes an m. Read every number against the scan. Search the Word file for characters that should not be there, such as a pipe or a tilde, which mark rules and smudges the engine tried to read as text.

A faint scan, a skewed page or a photograph with a shadow across it produce far more errors. If the source paper is still to hand, Scan documents with your phone covers getting a flat, well-lit capture that OCR can read.

Layout after OCR

OCR delivers words and, at best, paragraphs. Columns, tables and headings come back as plain text in reading order, and a table becomes lines with spaces between values. Rebuild the structure in Word with styles and Insert Table; the techniques in Keep formatting when converting PDF to Word apply, with the difference that here there was no structure to preserve.

PDF to Word OCR free of charge exists in every route above except some desktop programs, and every one of them costs proofreading time instead. The remaining Word guides are under PDF, Word, Excel and PowerPoint help.

Questions people ask

How do I know if my PDF is a scan?
Try to select a word. Text highlights word by word. A scan highlights as one block or not at all, and searching for a word finds nothing.
Why does PDF to Word report zero text for my file?
The PDF has no text layer. Each page is a picture, so there is nothing to convert. Run an OCR step first, then convert the searchable PDF it writes.
Can Word OCR a scanned PDF?
No. Word converts text layers only. A scan opens as pictures. OCR has to run first, in Preview on a Mac, in a desktop or scanner app, or through Google Docs.
Does OCR keep the layout of the scanned page?
Rarely. It returns words in reading order. Columns, tables and headings are rebuilt by hand in Word afterwards.
Is OCR through Google Docs private?
No. The PDF is sent to Google to be read. For a confidential document use Preview's Live Text on a Mac or a local OCR app on Windows.
How accurate is OCR on a scan?
High on a clean, straight 300 dpi scan and poor on a faint or skewed one. Errors concentrate in digits and names, so proofread those against the original.

Do it now

The tool runs in this browser. Your file never leaves the machine, and the result is checked before you download it.

PDF to Word

Where to go next