Skip to content
TNToolsNexus

Scanned PDF to text (OCR)

Extract text from a scanned PDF with OCR — each page is rendered and read in your browser, then the text is yours to copy. Free, private, nothing uploaded.

Drop a scanned PDF, or choose one

Each page is rendered and read on your device. Never uploaded.

Text recognition runs entirely in your browser — your file is never uploaded.

Turn a scanned PDF into copyable text. Each page is rendered and read with OCR right in your browser, so a document that was just images of pages becomes text you can edit and search — without uploading anything.

How to OCR a scanned PDF

  1. Drop or choose the scanned PDF.
  2. Click Extract text — the progress shows each page being read on your device.
  3. Copy the combined text.

Because OCR runs once per page, a long scan takes a little while; the first run also downloads the recognition engine (a few MB), then caches it.

Scanned vs. text PDFs

If your PDF already has selectable text, you don’t need OCR — the plain PDF to text tool pulls it out instantly. This tool is for scans and photos saved as PDF, where the pages are images: it renders each page and recognizes the words visually. If you’re not sure which you have, try selecting text in your PDF reader — if you can’t, it’s a scan, and this is the tool.

Private by design

Page rendering and text recognition both run locally in your browser (pdf.js + Tesseract, compiled to WebAssembly). Your scanned document is never uploaded — which is what you want for records and paperwork. For a single image rather than a PDF, use image to text.

Last updated:

Frequently asked questions

What's the difference between this and a normal PDF text extractor?
A normal extractor reads text that's already embedded in the PDF — it fails on scans, which are just images of pages. This tool uses OCR: it renders each page to an image and recognizes the text visually, so it works on scanned or photographed documents. If your PDF already has selectable text, the plain PDF-to-text tool is faster.
Is my PDF uploaded anywhere?
No. Each page is rendered and recognized entirely in your browser using WebAssembly — your document is never sent to a server. For scanned contracts, records, or personal paperwork, that local-only processing is the point.
How long does it take?
OCR is real computation, and it runs once per page, so a multi-page scan takes noticeably longer than reading a single image — expect a few seconds per page. A progress indicator shows which page is being read. The first run also downloads the recognition engine (a few MB), then caches it.
The text came out messy — why?
Scan quality is almost everything. Low-resolution scans, skewed pages, faint print, or heavy background noise all hurt accuracy. A cleaner scan (300 DPI, straight, good contrast) reads far better. OCR also does not read handwriting.
Does it keep the original layout?
It extracts the text in reading order, page by page, but it does not reproduce columns, tables, or exact formatting — you get the words, not a visual copy. For a faithful layout you'd keep the PDF; this is for getting the text out to edit or search.

Related tools

Image to Text (OCR)

Extract text from an image (JPG, PNG, screenshot, photo) with OCR that runs entirely in your browser — copy the recognized text. Free, private, no upload.

PDF to Text

Extract all text from a PDF into plain text you can copy or download — instant, free, and the file never leaves your browser.

Compress PDF

Shrink PDF file size by recompressing embedded photos and scans — text stays sharp and selectable. Free, private, no upload.