Skip to content
AsliPDF

OCR PDF

Make scanned PDFs searchable and selectable.

Runs on this device Free · no signup

How to use OCR PDF

  1. 01

    Drop in a scanned PDF or image

    Add the document whose pages are pictures of text rather than real, selectable text.

  2. 02

    Choose the language

    Pick the document's language so the recogniser uses the right model; the model downloads once on first use and is then cached.

  3. 03

    Download the searchable PDF

    Get back a PDF that looks identical but now has an invisible, selectable and searchable text layer behind the image.

Tested and last verified on

About OCR PDF

Run optical character recognition over scanned documents to add a searchable, selectable text layer and extract the recognized text — in multiple languages.

Make a scanned document searchable and selectable

A scanned PDF is really a stack of photographs: to your eyes it's a document, but to a computer each page is just an image with no idea what words it contains. That's why you can't select a sentence, why searching finds nothing, and why copying produces nothing useful. Optical character recognition bridges that gap. It looks at the picture, works out what the letters and words are, and adds an invisible text layer precisely aligned behind the image — so the page still looks exactly as scanned, but now behaves like a real text document that you can search, select, copy and extract from.

This turns a pile of scans into genuinely usable documents. A searchable archive of receipts, a contract you can quote from, a report whose figures you can pull out — all of it becomes possible once the text underneath is machine-readable. It also unlocks the rest of the toolkit, because tools that work on text, from extraction to redaction, need a real text layer to act on.

Getting the best results, privately

Recognition quality depends on a few things you can influence. Choosing the correct language matters, because the recogniser uses a language-specific model to decide between similar-looking characters, and the right model makes a clear difference for accented and non-Latin text. Clean input helps too: a straight, high-contrast scan is far easier to read than a skewed, faint or noisy one, so running a crooked scan through the Deskew tool first, or capturing a cleaner image to begin with, improves the outcome. Reviewable confidence means you can tell where the recogniser was unsure rather than trusting a hidden, possibly wrong, text layer.

Everything happens in your browser: the language model downloads once and is cached, and from then on the recognition itself runs locally with no upload, no account and no watermark. That's important because the documents people most often need to make searchable — identity papers, medical records, financial statements — are exactly the ones that should never be sent to someone else's server just to be read.

Questions

Everything runs in your browser. How we keep files private

What does OCR actually do to my PDF?

It reads the text in a scanned image and adds an invisible text layer aligned behind the picture. The page looks exactly the same, but you can now select, copy and search the words, and other tools can extract them.

Which languages are supported?

You choose the document's language before running, so the recogniser loads the matching model. Picking the correct language noticeably improves accuracy, especially for accented or non-English text.

Why does it need to download something the first time?

The language model is fetched once on first use and then cached in your browser. That initial download is the only time an internet connection is required; the recognition itself runs locally.

Is my scanned document uploaded to be read?

No. Recognition runs in your browser using WebAssembly, so a scanned contract, ID or record is processed on your own device.