PDF tool / 11
Scanned PDF to Text (OCR)
Read the text out of a scanned PDF that has no selectable text, page by page, entirely in your browser.
Scanned PDF to Text (OCR): TOEA renders every page to an image, runs optical character recognition over each one, and labels the output by page so you can tell where each came from. Runs 100% locally in your browser with zero server file uploads.
- Category
- PDF tools
- Runs
- In your browser
- Cost
- Free · no sign-up
- Availability
- Ready to use
Runs entirely in your browser
PDF · up to 30 MB, never uploaded
Getting a cleaner read
Every page is rendered at 200 DPI before it is read, so the limit is the scan itself. Body text scanned at 300 DPI reads well; a phone photo of a page taken at an angle, or a fax-quality scan of small print, produces guesses. Pages that are sideways or upside down read very badly, so turn them upright with rotate PDF before running the recognition.
Pick the language the document is written in, since each one is a separate model. Chinese comes in Simplified and Traditional models, and choosing the wrong one costs a lot of accuracy. A document mixing two languages reads best under the one used for most of the text.
Checking what came out
The confidence figure is an average across all pages, so one bad page can hide inside a good score. Around 90 or above usually means only occasional slips; below 75, expect errors in most paragraphs. The usual confusions are 0 and O, 1, l, and I, rn read as m, and 5 as S, which matters most in amounts, dates, and reference numbers, so check every figure against the page.
Leave Join lines into paragraphs on for prose, where it mends sentences broken at the end of each printed line. Turn it off for addresses, lists, and tables, where the line breaks carry meaning.
How to use it
- Choose a scanned PDF.
- Pick the language of the text.
- Read the result and copy or download it.
Privacy & limitations
Your file is never uploaded. The recognition engine and language model are downloaded once — around 10 MB — and cached by your browser, then everything runs on your own machine. Hosted OCR services work the other way round: they want the document. A long document takes a while, because each page is rendered and then read in turn.
Related tools
Frequently asked questions
When do I need this rather than PDF to Text?
Use PDF to Text first — it is instant and exact when the document has real text inside it. This tool is for the case where that returns nothing, which means the pages are photographs.
Why is it slow?
Every page is rendered to an image and then read character by character. A few pages is quick; a hundred-page scan takes real time, and the tab has to stay open.
Does it keep the layout?
No. Columns, tables, and headers come out as plain lines in reading order, so a complex page will need tidying.
Can it make the PDF searchable?
Not here — that means writing an invisible text layer back into the file. This gives you the text to use elsewhere.
Free tool · runs in your browser · no account required