OCR PDF

Make a scanned PDF searchable — the pages look the same, but their text can be found, selected and copied.

Drop your scanned PDF here
or click to browse — the file stays on your device

Pick the language actually printed in the document — a wrong one returns garbled text. Add a second language only when the pages really mix two, as each extra one slows recognition down.

Pages that already contain real text are normally left alone. Tick the box to recognise them anyway, for example a scan with a typed header; on pages that were searchable already this can double up the text.

The first run downloads the recognition engine and the language data into your browser cache — a few megabytes, once. Your PDF itself is never uploaded.

Your searchable PDF is ready

How to make a scanned PDF searchable

  1. Drop the scanned PDF onto the box above, or click it to browse. One document is processed at a time; if it is locked with a password, you are asked for it first.
  2. Set Text language to the language the document is written in. For pages that genuinely mix two, such as a bilingual contract, choose a Second language too; otherwise leave it on None.
  3. Press Make PDF searchable. Every page is checked first: pages that already hold real text are kept as they are, and the rest are drawn at up to 300 DPI and read one at a time. The status line names the current page, and Cancel stops the job at any point.
  4. The finished copy downloads as yourfile-searchable.pdf. If all you need are the words, Download text only (.txt) saves them as a plain UTF-8 file with a marker at the start of each page.

What people use it for

A scan is a photograph of paper. To your computer it is a picture with no words in it: search finds nothing, a sentence cannot be copied out, and the file never turns up when you type a phrase into your file manager. OCR adds the missing words behind the picture without changing how the pages look.

Good to know

The quality of the result comes down to the scan and the language setting. Straight, sharp, evenly lit pages read well. Faint photocopies, curled or skewed pages, tiny print and handwriting produce mistakes, and the wrong language produces nonsense. Each recognised word sits invisibly over the spot where it appears, so selecting text highlights the right area; a misread word, however, is one search will not find. The recogniser expects pages the right way up, so put sideways scans through Rotate PDF first. Your original pages are kept exactly as they were, with only a thin text layer added, so the file usually grows only a little.

All of the work happens on your own device. The first time you use a language, the recognition engine and that language's data are downloaded into your browser cache: code and dictionaries travelling to you, nothing of yours travelling out. The PDF is opened, drawn and read inside this tab, and the searchable copy goes straight to your downloads. That matters more for OCR than for most jobs, because the papers people scan tend to be the sensitive ones: signed agreements, medical letters, tax forms and copies of ID.

Questions people ask

Will my PDF look any different?

No. The original pages stay as they were, at their original resolution, and the added text is invisible. You only notice it when you search, select or copy.

Why does a long document take a while?

Every page without text is analysed on your own device, which takes a noticeable moment per page, and longer on a phone than on a computer. Pages that already contain text are skipped, and only one page image is kept in memory at a time.

Can I get just the text?

Yes. After a run, Download text only (.txt) saves everything that was read. For a single photo or screenshot, Image to Text is quicker, and for a PDF that already has real text, PDF to Text extracts it exactly, with no recognition step to get wrong.

Why were some of my pages skipped?

A page that already contains more than a few characters of real text counts as searchable and is left untouched, so a file that mixes typed pages with scans only has its scans recognised. If a scanned page carries a typed header or stamp, tick Recognise pages that already have text to process it anyway.