How to Extract Text from a Scanned PDF or Image
A scanned document is just a picture of words — you can't select, copy, or search it. OCR fixes that. Here is how to pull real text out of a scan or photo, privately.
You’ve got a scanned contract, a photographed receipt, or a PDF someone made by scanning paper — and you need to copy a few lines out of it. But nothing selects. That’s because a scan is an image of text, not text. The fix is OCR.
Why you can’t select text in a scan
When a document is scanned or photographed, the result is a picture. The letters are just arrangements of pixels — the computer has no idea they’re words. So you can’t select, copy, search, or edit them. A PDF can contain either real text or scanned images (or both), which is why some PDFs let you select text and others don’t.
OCR turns the picture back into text
Optical Character Recognition (OCR) looks at the shapes in the image and recognizes them as letters and words, producing real, selectable text. Two tools here do it, depending on what you have:
- A photo or screenshot of text → the image-to-text tool.
- A scanned PDF → the scanned-PDF-to-text tool, which reads each page.
If your PDF was born digital (exported from Word, say) rather than scanned, you don’t need OCR at all — the PDF-to-text extractor pulls its existing text directly, which is faster and perfectly accurate.
Scan vs. digital: how to tell
Open the PDF and try to select a line. If the cursor selects text, it’s digital — use the straightforward text extractor. If it selects nothing (or selects the whole page as one block/image), it’s a scan — use OCR.
Getting good OCR results
- Straight and well-lit beats crooked and shadowed. A flat, evenly-lit scan reads far better than a photo taken at an angle.
- Higher resolution helps up to a point — 300 DPI scans are the sweet spot.
- Clean printed text OCRs almost perfectly; handwriting and stylized fonts are much harder, so treat those results as a rough draft.
The privacy part
OCR is exactly the kind of thing you shouldn’t upload — scans are often IDs, contracts, medical forms, and financial records. These tools run the recognition entirely in your browser: the image is read on your device and never sent to a server. Do it locally, and a sensitive document never leaves your machine just because you needed to copy a paragraph out of it.
Last updated: