Guides & Tutorials

OCR: how to recognise the text of a scanned document

OCR turns the picture of a page into real text: selectable, searchable and editable. Here the recognition runs entirely on your own computer, without sending the scan to any service.

1. What OCR is

OCR stands for Optical Character Recognition: it is the technology that looks at the image of a page and extracts the text from it, letter by letter.

It is needed because, to a computer, a scan is not a document but a photograph. The words you read perfectly well are, for software, merely light and shade: they cannot be selected, searched, copied or corrected.

OCR bridges exactly that gap. After recognition the page contains real text and the document behaves like a normal PDF again.

🔒
In 123 PDF Editor PRO OCR runs on an engine built into the browser: the scan is not uploaded to any server, not even during recognition. That is a substantial difference from online OCR services, where the document — often a deed, an invoice or an identity card — is sent to a third-party infrastructure.

🔍 2. When you actually need it

Not every PDF needs OCR. There is a simple way to tell in two seconds.

Try selecting a word with the mouse. If the text highlights, the document already contains real text and OCR is pointless. If instead the cursor makes you draw a rectangle as it would on a photo, you are looking at a scan.

  • Paper documents run through a scanner or photographed with a phone.
  • Faxes, receipts and forms archived years ago in image format.
  • PDFs produced by older business systems that print to an image.
  • Screenshots and photos of pages inserted into a PDF.

3. Running OCR

The editor spots pages without text by itself and offers recognition, but you can also start it whenever you like.

The editor detects scanned pages and offers OCR
Pages without selectable text are detected automatically.
  1. Install the extension from the download page; it is a Chrome extension and works the same on Windows, Mac and Linux.
  2. Open the document with Choose a PDF (or Open PDF).
  3. If the Scanned pages detected window appears, carry on from there; otherwise open Local OCR from the toolbar.
  4. Choose the language of the document, or leave Automatic.
  5. Click Run local OCR and wait for the progress bar to complete.
  6. When it finishes the text is selectable: you can correct it, search it and copy it.
The text of a scanned PDF becomes editable after OCR
After recognition the flat scan turns into editable, searchable text.

🌐 4. The languages recognised

10 languages are included, all available offline: nothing has to be downloaded at the moment of use.

  1. If you know the language of the document, select it: recognition is faster and more accurate.
  2. If the document is mixed or you are unsure, leave Automatic: all the included languages are loaded.
  3. The automatic option takes more time and more memory, so it is worth using only when you really need it.
💡
Language mostly affects accented letters and characters that look alike. Telling the engine it is reading English on an English document avoids a good share of the errors.

5. Getting a good result

The quality of OCR depends almost entirely on the quality of the original scan: no engine recovers what is not in the image.

  • Adequate resolution — a scan at 300 DPI is the usual reference for text; below 200 DPI errors increase sharply.
  • A straight page — if the scan is crooked, straighten it first with page rotation.
  • Good contrast — black text on white works better than a faded photocopy or a photo taken in poor light.
  • Printed text — OCR reads typeset characters; handwriting is out of reach.
  • Always proofread — on figures, reference codes and amounts a manual check pays off: a 0 read as O slips by easily.

6. What you can do after recognition

Once the text exists, the document behaves like any other PDF.

  1. Correct and update the text: see the guide on how to edit the text of a PDF.
  2. Search for a word across the whole document, which is impossible on a scan.
  3. Highlight the important passages, as explained in the guide on how to highlight a PDF.
  4. Remove confidential data properly, following the guide on how to redact text.
  5. Archive the document as a searchable file, so you can find it later by searching for a word.

? 7. Frequently asked questions

The questions that come up most often on this topic.

Does OCR work without an internet connection?

Yes. The recognition engine is bundled with the extension and works locally: you can run OCR completely offline.

How long does it take?

It depends on the number of pages, the resolution and your computer. A few pages is a matter of seconds; a long document scanned at high resolution takes more patience, especially with the automatic option across all languages.

Is the recognised text identical to the original?

Almost always very close, but no OCR is infallible. On important documents, at least proofread the numbers and proper names.

Does the free version put a watermark on OCR exports?

OCR exports in the free version carry a small “Free Version” marking. The PRO licence, a one-time payment of $24, removes it along with the limits on downloads and projects.

Make your scan readable

Install the extension, open the document and run recognition — without sending the scan online.